Cardiovascular

Machine learning and deep learning-based prediction of hypertension and analysis of its major risk factors in Bangladesh.

TL;DR

Hypertension prevalence in Bangladesh was 18.04% and, among machine learning and deep learning models, weighted logistic regression achieved the highest accuracy and specificity while random forest achieved the highest recall and F1-score, with age, BMI, sex, family size, and educational level identified as the most important predictors.

Key Findings

The overall prevalence of hypertension among adults in Bangladesh was 18.04%, with higher prevalence among women than men.

  • Overall prevalence was 18.04% (95% CI: 17.2%–18.9%)
  • Prevalence was higher among women (18.87%) than men (16.97%)
  • Data were drawn from 14,283 adults aged 18 years or older from the 2022 Bangladesh Demographic and Health Survey
  • The study used a cross-sectional design

Hypertension was significantly associated with age, BMI, diabetes, wealth index, education, household size, and region.

  • Chi-square tests were used to assess associations between hypertension and socio-demographic variables
  • All listed variables showed statistically significant associations (p < 0.05)
  • Variables assessed included age, BMI, diabetes, wealth index, education, household size, and region

Weighted logistic regression (WLR) achieved the highest accuracy, precision, specificity, AUC-ROC, and AUC-PR among all models tested.

  • WLR accuracy: 0.817, precision: 0.444, specificity: 0.981, AUC-ROC: 0.751, AUC-PR: 0.357
  • However, WLR exhibited low recall (0.070), limiting its utility for identifying individuals with hypertension
  • Four ML models were compared: weighted logistic regression, random forest, extreme gradient boosting, and light gradient boosting machine
  • Two DL models were also compared: TabNet and multi-layer perceptron

The random forest (RF) model achieved the highest recall and F1-score on the test data, indicating greater sensitivity in identifying individuals with hypertension.

  • RF achieved the highest recall (0.687) and F1-score (0.460) on the test data
  • The authors indicate RF may be more suitable for public health applications because of its higher recall and F1-score
  • The authors note that further external validation and assessment of clinical utility are required before implementation

Age, BMI, sex, family size, and educational level were identified as the most important predictors of hypertension among the variables included in the study.

  • Variable importance was assessed across the machine learning and deep learning models
  • These five variables were identified as the top predictors from among all socio-demographic and health-related variables included
  • Diabetes, wealth index, and region were also significantly associated with hypertension in chi-square analyses but ranked lower in model-based feature importance

Model performance was evaluated using multiple metrics including accuracy, precision, recall, specificity, F1 score, AUC-ROC, and AUC-PR to account for the imbalanced nature of the dataset.

  • The dataset was imbalanced given the 18.04% hypertension prevalence, motivating use of weighted logistic regression and evaluation via AUC-PR in addition to AUC-ROC
  • Six models in total were applied and compared across all metrics
  • The use of both AUC-ROC and AUC-PR was noted as important for imbalanced classification tasks

What This Means

This research suggests that nearly 1 in 5 adults in Bangladesh (about 18%) has hypertension, with women slightly more affected than men. Using data from over 14,000 adults surveyed in 2022, the researchers found that factors such as older age, higher BMI, having diabetes, lower education, household size, wealth, and geographic region were all meaningfully linked to hypertension risk. Among these, age, BMI, sex, family size, and education level stood out as the strongest predictors when fed into computer-based prediction models. The study tested six different artificial intelligence models — four machine learning and two deep learning approaches — to see which could best identify people with hypertension. No single model excelled at everything. The weighted logistic regression model was the most accurate overall and rarely gave false alarms, but it missed most actual hypertension cases (catching only about 7%). The random forest model, by contrast, correctly identified about 69% of people who actually had hypertension, making it potentially more useful for public health screening where missing cases is costly. The authors caution that before any model is used in practice, it would need to be tested on new, independent datasets to confirm its reliability. This research matters because hypertension is a major driver of heart disease and stroke, and Bangladesh, like many low- and middle-income countries, faces growing burdens from these conditions. The findings highlight that socioeconomic and demographic factors — not just biology — play a central role in hypertension risk, pointing to the importance of targeted public health programs. The comparison of prediction models also offers practical guidance for researchers and health planners on choosing the right tool depending on whether minimizing missed cases or minimizing false positives is the priority.

Have a question about this study?

Citation

Chandra S, Molla M, Islam S, Rahaman M, Ali M, Ali M. (2026). Machine learning and deep learning-based prediction of hypertension and analysis of its major risk factors in Bangladesh.. PloS one. https://doi.org/10.1371/journal.pone.0358471