An interpretable XGBoost model based on six clinical, laboratory, and CGM-derived indicators showed acceptable discrimination, calibration, and clinical net benefit for estimating ultrasound-defined plaque vulnerability among patients with T2DM and established carotid plaque.
Key Findings
Results
Six core predictors of carotid plaque vulnerability were identified in T2DM patients: systolic blood pressure (SBP), LDL-C, age, time in range (TIR), systemic immune-inflammation index (SII), and smoking status.
Core predictors were identified by integrating feature-importance rankings derived from three algorithms: Random Forest, linear Support Vector Machine, and Logistic Regression.
These six predictors were selected from demographic, clinical, biochemical, continuous glucose monitoring (CGM), and inflammatory indicators collected at two medical centers.
CGM was performed for 72 consecutive hours using a retrospective CGM system (Medtronic iPro2), and TIR was calculated from valid CGM recordings.
Results
The XGBoost model demonstrated the best overall validation performance among the five predictive models evaluated.
The XGBoost model attained an AUC of 0.882 (95% CI: 0.795–0.969) in the external validation set.
Five predictive models were developed and evaluated using discrimination, calibration, and clinical-utility metrics.
Model evaluation included external validation in addition to internal validation to assess generalizability.
Results
SHAP analysis revealed that higher SBP, higher LDL-C, older age, higher SII, and smoking were each associated with a higher predicted probability of vulnerable plaques.
SHAP (SHapley Additive exPlanations) analysis was used to provide interpretability for the XGBoost model.
These five factors were positively associated with predicted plaque vulnerability probability.
The direction and magnitude of each predictor's contribution to individual predictions was made explicit through SHAP values.
Results
Higher time in range (TIR) derived from continuous glucose monitoring was associated with a lower predicted probability of vulnerable carotid plaques.
TIR was calculated from 72 consecutive hours of CGM data using the Medtronic iPro2 retrospective system.
TIR was the only CGM-derived metric among the six core predictors.
SHAP analysis indicated that higher TIR was associated with a lower predicted probability of vulnerable plaques, in contrast to the other five predictors which were positively associated.
Methods
The study enrolled 884 T2DM patients with carotid atherosclerotic plaques from two medical centers, allocated into training, internal validation, and external validation cohorts.
Data were collected retrospectively from two medical centers.
Data preprocessing included clinically reviewed missing-data handling, standardization of continuous variables, and Synthetic Minority Over-sampling Technique (SMOTE).
SMOTE was fitted only within the training data and, during cross-validation, within each training fold to avoid information leakage.
Plaque vulnerability was defined by ultrasound criteria.
Conclusions
The authors concluded that the model is not a substitute for carotid ultrasound and may support research-stage prioritization for expert plaque characterization when imaging capacity or expertise is constrained.
The authors explicitly stated: 'The model is not a substitute for carotid ultrasound.'
Pending prospective workflow and economic evaluation, the model may support research-stage prioritization for expert plaque characterization.
The model's clinical application is framed as potentially useful when imaging capacity or expertise is constrained.
The framework was described as 'interpretable,' with SHAP analysis used to explain individual predictions.
What This Means
This research suggests that a type of machine learning model called XGBoost can predict which diabetic patients with carotid artery (neck artery) plaques are at higher risk of having 'vulnerable' plaques—plaques that are more likely to rupture and cause strokes or other serious events. The model was built using data from 884 patients at two hospitals and was tested on patients from a separate hospital to confirm it worked beyond the original dataset. It achieved strong predictive accuracy (AUC of 0.882 in external testing), using just six pieces of information: blood pressure, LDL cholesterol, age, a measure of blood glucose control called 'time in range' (TIR) from a continuous glucose monitor, an inflammation score (SII), and smoking status.
A key finding is that better glucose control—specifically, spending more time with blood sugar in a healthy range as measured by a continuous glucose monitor worn for 72 hours—was associated with a lower predicted risk of having a dangerous plaque. Higher blood pressure, higher LDL cholesterol, older age, more inflammation, and smoking were each linked to higher predicted risk. The model also used 'explainability' tools (called SHAP analysis) so that clinicians can see exactly why the model made a particular prediction for a specific patient, rather than treating it as a 'black box.'
This research suggests that combining routine clinical measurements with short-term continuous glucose monitoring data could help identify which type 2 diabetes patients most urgently need detailed carotid ultrasound evaluation, particularly in settings where specialist imaging is limited. The authors are careful to note that this model is not meant to replace ultrasound imaging, and that further prospective studies are needed before it could be used in routine clinical care.
Cheng Q, Zhang F, Wang H, Qian Y, Tang P, Cai W, et al.. (2026). Prediction of carotid vulnerable plaques in patients with type 2 diabetes using interpretable machine learning models.. Frontiers in endocrinology. https://doi.org/10.3389/fendo.2026.1911126