Machine learning models using routinely available clinical and radiographic variables did not achieve clinically actionable risk stratification for cSDH recurrence, with discriminative capacity insufficient to identify a low-risk subgroup suitable for de-escalated surveillance.
Key Findings
Results
Postoperative recurrence requiring reoperation occurred in 30.1% of patients in this cohort.
170 of 564 consecutive patients experienced postoperative recurrence requiring reoperation.
The study included patients undergoing surgical cSDH evacuation from 2015 to 2023 at a single center.
This recurrence rate is consistent with the reported literature range of 5-33%.
The dataset was randomly divided into training (75%, n=422) and test (25%, n=142) sets.
Results
XGBoost achieved the highest cross-validated ROC AUC in the training set, outperforming logistic regression but not Random Forest.
XGBoost achieved a cross-validated ROC AUC of 0.713 (SE=0.024) on the training set.
Random Forest matched XGBoost performance on the training set.
Logistic regression achieved a lower cross-validated ROC AUC of 0.686.
Model development and tuning used 10-fold cross-validation on the training set.
Three models were compared: regularized logistic regression, Random Forest, and XGBoost, using 31 predictor variables.
Results
The final XGBoost model achieved a ROC AUC of 0.688 on the held-out test set, indicating modest discriminative performance.
Test set ROC AUC was 0.688 (95% CI 0.590-0.772).
The model showed satisfactory average calibration but overconfident individual-level risk estimates.
The calibration slope was 0.615, indicating overconfidence in individual-level predictions.
Consistency between training and test performance indicated limitations reflect predictor information content rather than overfitting.
Results
Hematoma volume, coagulation parameters, and disease severity markers were the most influential predictors, though effect sizes remained modest.
The most influential predictors included hematoma volume, coagulation parameters, ICU admission, and Glasgow Coma Scale (GCS) score.
Despite being the top predictors, effect sizes across all variables remained modest.
All 31 predictor variables were derived from routinely available clinical and radiographic data.
The modest effect sizes of even the top predictors contributed to the overall limited discriminative performance.
Results
At a clinically relevant 90% sensitivity threshold, specificity was only 30.3%, allowing potential imaging reduction in roughly one-third of non-recurrence patients.
At 90% sensitivity, the model achieved a specificity of only 30.3%.
This threshold would potentially allow imaging reduction in roughly one-third of non-recurrence patients.
The authors characterized this level of specificity as insufficient to identify a low-risk subgroup suitable for de-escalated surveillance.
The 90% sensitivity threshold was described as 'clinically relevant' in the context of not missing true recurrence cases.
Discussion
The study concluded that cSDH recurrence is likely driven by factors not captured in standard clinical assessment, supporting uniform or symptom-driven imaging strategies over risk-stratified approaches.
The authors state that 'recurrence is driven by factors not captured in standard clinical assessment.'
The findings do not support risk-stratified surveillance based on routinely available variables.
The authors recommend 'uniform or symptom-driven imaging strategies over risk-stratified approaches.'
Predicting recurrence was proposed as a means to enable risk-stratified surveillance, reducing imaging in low-risk patients while maintaining monitoring for high-risk individuals — a goal this study found unachievable with current variables.
Methods
This was a retrospective, single-center study with internal validation only, limiting the generalizability of the findings.
The study included 564 consecutive patients from a single center over an 8-year period (2015-2023).
Only internal validation was performed using a held-out test set; no external validation cohort was used.
The authors acknowledge this as a limitation when interpreting the results.
The single-center design may limit applicability to other institutions or patient populations.
What This Means
This research examined whether artificial intelligence (machine learning) could predict which patients would need a second surgery after treatment for chronic subdural hematoma (cSDH) — a type of blood clot that collects on the surface of the brain. Using data from 564 patients treated at a single hospital over eight years, the researchers tested three different machine learning approaches using 31 variables routinely collected during clinical care, such as blood clot size, blood clotting test results, and measures of how sick the patient was at admission. About 30% of patients in this study experienced a recurrence requiring reoperation, which is consistent with rates reported in prior research.
The best-performing model (XGBoost) was only modestly accurate at distinguishing patients who would need reoperation from those who would not, achieving an AUC of about 0.69 on an independent test group — a value where 1.0 would be perfect and 0.5 would be no better than chance. When the model was set to catch 90% of true recurrences (high sensitivity), it could only correctly rule out recurrence in about 30% of patients who did not actually recur (low specificity). This means the model could not reliably identify a 'low-risk' group that could safely skip follow-up imaging. The researchers noted that performance limitations appeared to reflect the information content of the available variables rather than technical issues with the models themselves.
This research suggests that the factors driving cSDH recurrence may not be well captured by standard clinical measurements currently recorded in patient records, making accurate individual-level prediction very difficult with existing data. Rather than supporting risk-based imaging schedules — where low-risk patients get less monitoring — the findings suggest that uniform follow-up or symptom-driven imaging strategies may be more appropriate until better predictive variables or external validation studies are available. Future work may need to identify new biological or imaging markers associated with recurrence risk.
Hamou H, Kernbach J, Ridwan H, Fay-Rodrian K, Clusmann H, Hoellig A, et al.. (2026). Classification of recurrence status after surgical treatment of chronic subdural hemorrhage - A machine learning approach.. PloS one. https://doi.org/10.1371/journal.pone.0346756