Cardiovascular

Development and temporal validation of a machine learning-based prediction model for long-term depressive symptoms in older adults with cardiovascular disease or hypertension: A longitudinal cohort study.

TL;DR

A logistic regression model using seven predictors demonstrated good discrimination and calibration for predicting long-term depressive symptoms among older adults with cardiovascular disease or hypertension, achieving an AUC of 0.844 in development and 0.840 in temporal validation cohorts.

Key Findings

Logistic regression outperformed six other machine learning algorithms in predicting long-term depressive symptoms in older adults with cardiovascular disease or hypertension.

  • Seven machine learning algorithms were compared in the model development phase.
  • Logistic regression showed the best overall performance among all algorithms tested.
  • The development cohort included 548 participants followed from 2011 to 2015.
  • The temporal validation cohort included 523 nonoverlapping participants followed from 2015 to 2020.
  • Model selection was based on overall performance metrics including discrimination and calibration.

The logistic regression model achieved an AUC of 0.844 in the development cohort and 0.840 in the temporal validation cohort.

  • Development cohort AUC: 0.844 (95% CI, 0.773–0.915).
  • Temporal validation cohort AUC: 0.840 (95% CI, 0.799–0.881).
  • The narrow confidence interval in the validation cohort reflects the larger effective sample contributing to that estimate.
  • The model demonstrated good discrimination and calibration in both cohorts.
  • Temporal validation used a nonoverlapping cohort to assess generalizability across time periods.

The Boruta algorithm selected seven predictors for inclusion in the final prediction model.

  • The Boruta algorithm was used specifically for predictor selection from the available variable set.
  • Seven predictors were ultimately retained for the logistic regression model.
  • Data were drawn from the China Health and Retirement Longitudinal Study (CHARLS) spanning 2011 to 2020.
  • Missing values in the dataset were imputed using a random forest algorithm prior to model development.

A time-based split design was used to develop and temporally validate the prediction model using data from the China Health and Retirement Longitudinal Study.

  • Data were obtained from CHARLS covering the years 2011 to 2020.
  • The development cohort followed participants from 2011 to 2015 (n = 548).
  • The temporal validation cohort followed a nonoverlapping group of participants from 2015 to 2020 (n = 523).
  • The time-based design was chosen to simulate prospective model performance and assess temporal generalizability.
  • The study used a longitudinal cohort design.

Shapley additive explanations (SHAP) were used to interpret the contributions of individual predictors in the selected model.

  • SHAP values were applied to the final logistic regression model to provide interpretability.
  • This approach allows quantification of each predictor's contribution to individual predictions.
  • Use of SHAP was intended to support clinical interpretability of the machine learning-based model.
  • The authors noted that further independent and prospective validation is required before routine clinical implementation.

Older adults with cardiovascular disease or hypertension are identified as a population at elevated risk for depressive symptoms that may adversely affect prognosis and quality of life.

  • The study population was restricted to older adults diagnosed with cardiovascular disease or hypertension.
  • Depressive symptoms in this population were characterized as potentially adversely affecting prognosis and quality of life.
  • The authors identified a need for a practical prediction model to support early risk stratification and targeted screening in this group.
  • The long-term nature of the outcome (depressive symptoms assessed over multi-year follow-up) was a key study feature.

What This Means

This research suggests that a relatively straightforward statistical method called logistic regression, when informed by machine learning-based variable selection, can accurately predict which older adults with heart disease or high blood pressure will develop lasting depressive symptoms over a period of several years. The researchers used data from a large ongoing Chinese study tracking the health of older adults between 2011 and 2020, dividing the data into an earlier group to build the model and a later, separate group to test whether it still worked well over time. The model identified seven key risk factors and correctly classified patients with about 84% accuracy in both groups, suggesting it holds up well across different time periods. This matters because depression is common among people with heart disease and high blood pressure, and it can make those conditions worse while reducing quality of life. Having a tool that can flag who is at higher risk years in advance could help doctors and healthcare systems focus attention and resources on the people who need it most, potentially enabling earlier intervention. The use of SHAP values also means clinicians could see which specific factors are driving any individual patient's risk, making the model more transparent and usable in practice. This research suggests the model has strong potential as a clinical screening tool, but the authors caution that it has only been tested within one Chinese longitudinal study and in one country's population. Additional independent testing in other settings and prospective real-world studies would be needed before it could be routinely used in clinical care.

Have a question about this study?

Citation

Zhao Y, Chen S, Pang C, Tang Y, Long S, Zheng H, et al.. (2026). Development and temporal validation of a machine learning-based prediction model for long-term depressive symptoms in older adults with cardiovascular disease or hypertension: A longitudinal cohort study.. Medicine. https://doi.org/10.1097/MD.0000000000050520