Cardiovascular

Machine learning models for risk prediction of diabetic retinopathy in Fujian eye study.

TL;DR

Among five machine learning models trained on population-based data from the Fujian Eye Study, the Support Vector Classifier achieved the highest performance (F1-score 92.83%, AUC 0.99) for diabetic retinopathy risk prediction, with history of diabetes, age, pulse pressure difference, near visual acuity of the left eye, and height identified as the top five predictors.

Key Findings

The Support Vector Classifier (SVC) model achieved the highest predictive performance among all five machine learning models tested for diabetic retinopathy risk prediction.

  • SVC achieved an F1-score of 92.83% and AUC of 0.99.
  • The five models compared were Logistic Regression (LR), K-Nearest Neighbors (KNN), Support Vector Classifier (SVC), Decision Tree (DT), and Random Forest (RF).
  • Models were trained and optimized via cross-validation and grid search.
  • Performance was evaluated using accuracy, precision, recall, F1-score, and AUC.

SHAP analysis identified five key predictors of diabetic retinopathy risk: history of diabetes, age, pulse pressure difference (PPG), near visual acuity of the left eye, and height.

  • Feature importance was examined using SHAP (Shapley Additive Explanations).
  • These predictors were validated through three additional methods: Factor Analysis (FA), Highly Variable Feature Selection (HVGS), and Spearman's rank correlation.
  • Cross-method comparison confirmed high feature stability across models, indicating robust predictor reproducibility.
  • Pulse pressure difference (PPG) was among the top five predictors, suggesting a vascular component to DR risk in this population.

The study dataset comprised 8,211 participants with 51 variables drawn from the population-based Fujian Eye Study.

  • Data were obtained from the Fujian Eye Study, a population-based dataset.
  • The dataset included 8,211 participants and 51 variables.
  • Data preprocessing was performed prior to model training.
  • The study aimed to identify top predictors and establish a robust data-driven framework for early DR screening.

The machine learning framework demonstrated robust predictor reproducibility across multiple feature selection methods.

  • Three unsupervised and nonparametric approaches were used to validate SHAP-identified features: Factor Analysis (FA), Highly Variable Feature Selection (HVGS), and Spearman's rank correlation.
  • Cross-method comparison confirmed high feature stability across models.
  • The consistency across methods supports the reliability of the identified risk factors for DR.

The authors conclude that the machine learning-based framework provides a reliable, data-driven tool for early DR screening that can facilitate population-level risk stratification.

  • The identified key risk factors are described as providing 'actionable insights for early stratification.'
  • The framework is proposed to 'inform personalized preventive interventions in clinical and public health settings.'
  • Near visual acuity of the left eye and height were identified as novel predictors not typically emphasized in traditional DR risk models.
  • The study positions the SVC model as suitable for population-level screening applications.

What This Means

This research suggests that machine learning models can accurately predict who is at risk for diabetic retinopathy (DR), a serious eye complication of diabetes that can lead to vision loss. Using data from over 8,000 people in the Fujian Eye Study in China, the researchers tested five different machine learning approaches and found that a method called the Support Vector Classifier performed best, correctly identifying DR risk with very high accuracy (AUC of 0.99 and F1-score of nearly 93%). The models were trained on 51 different health and demographic variables, and the researchers used multiple analytical techniques to determine which factors were most predictive. The study found that the five most important predictors of diabetic retinopathy were: having a history of diabetes, older age, a higher pulse pressure difference (the gap between systolic and diastolic blood pressure, which reflects vascular health), reduced near vision in the left eye, and height. Some of these factors, such as pulse pressure difference and height, are not typically emphasized in standard clinical assessments for DR risk. The fact that these predictors were consistently identified across multiple different analytical methods adds confidence that they are genuinely important rather than statistical artifacts. This research suggests that machine learning tools could be used in clinical and public health settings to screen large populations for DR risk more efficiently, potentially catching high-risk individuals earlier when interventions are more effective. Rather than relying solely on traditional risk factors like blood sugar control, these models incorporate a broader set of variables to provide more nuanced risk stratification. However, as with any model developed in one population, further validation in diverse populations would be important before widespread clinical adoption.

Have a question about this study?

Citation

Li Y, Hu Q, Wang B, Luo X, Zhang M, Li X. (2026). Machine learning models for risk prediction of diabetic retinopathy in Fujian eye study.. Frontiers in endocrinology. https://doi.org/10.3389/fendo.2026.1834380