XGBoost showed modestly higher discriminative performance than conventional logistic regression or the RACE scale for prehospital detection of anterior-circulation large vessel occlusion or intracranial haemorrhage, while conventional logistic regression showed better calibration, and absolute gains were limited.
Key Findings
Results
Using routine prehospital variables, supervised machine learning (XGBoost) achieved higher balanced accuracy than conventional logistic regression and the RACE scale for detecting aLVO or ICH.
Mean balanced accuracy was 82.1% for XGBoost (sML), 79.0% for conventional logistic regression (cLR), and 74.5% for the RACE scale
The difference between methods was statistically significant (P < .01)
Routine variables included demographics, symptom duration, vital signs, glucose, and neurological deficits
Analysis used 2-fold internal-external cross-validation across two datasets (2018–2019)
Results
Using an extended variable set, both machine learning and logistic regression showed improved performance compared to routine variables alone.
With extended variables, balanced accuracy improved to 84.4% for sML and 80.7% for cLR (P < .01)
The extended variable set additionally included medical history, anticoagulant use, and point-of-care international normalised ratio (INR)
The RACE scale was not evaluated with the extended variable set, as it is a fixed clinical scale
Results
The XGBoost model achieved the highest sensitivity and specificity among all compared methods.
sML achieved sensitivity of 76.7% and specificity of 92.1%
These represent the peak discriminative performance values reported in the study
These figures correspond to the extended variable set condition for sML
Results
Calibration favoured conventional logistic regression over XGBoost, as indicated by Brier scores.
Brier scores were lower (better) for cLR than for sML
Calibration was described as 'good for both models' based on calibration plots
Despite higher discrimination, XGBoost's calibration performance was inferior to cLR
Methods
The study population consisted of 3320 patients with emergency medical services-activated stroke codes, of whom 538 (16%) had the composite outcome of aLVO or ICH.
Data were drawn from two studies: the Leiden Prehospital Stroke Study (LPSS) and the Prehospital Triage of Patients With Suspected Stroke Study (PRESTO)
Patients were enrolled between 2018 and 2019
The composite outcome combined anterior-circulation large vessel occlusion (aLVO) and intracranial haemorrhage (ICH)
The prevalence of the composite outcome was 16% (538/3320)
Conclusions
The authors concluded that machine learning may improve prehospital classification but that absolute performance gains were limited and require further validation.
The paper states findings 'suggest that machine learning may improve prehospital classification in selected settings, but the absolute gains were limited'
The authors called for 'further external validation and prospective implementation studies'
Traditional triage based on clinical scales or logistic regression was noted to potentially miss complex predictor interactions
What This Means
This research compared three different methods for identifying stroke patients who need specialized emergency care before they reach the hospital. Specifically, the study looked at detecting strokes caused by large blocked blood vessels or bleeding in the brain — conditions requiring immediate treatment at specialized stroke centers. The three methods compared were: a machine learning algorithm called XGBoost, a traditional statistical approach called logistic regression, and an existing clinical scoring tool called the RACE scale. The study analyzed data from 3,320 patients who were evaluated by emergency medical services in the Netherlands between 2018 and 2019, of whom about 1 in 6 had the serious stroke types being targeted.
This research suggests that the machine learning approach (XGBoost) performed modestly better than the traditional statistical method and notably better than the clinical scale at correctly identifying patients with serious strokes. When extra information such as medical history and blood-thinning medication use was added, both the machine learning and statistical models improved further. However, the traditional statistical model (logistic regression) produced more reliable probability estimates — meaning its confidence levels were better calibrated to actual outcomes — even though it was slightly less accurate overall at distinguishing serious strokes from other conditions.
The practical implication is that machine learning tools could potentially help paramedics and emergency dispatchers make better decisions about where to send stroke patients, which matters because getting the right patient to the right hospital quickly can be lifesaving. However, the improvements over traditional methods were modest, and the authors caution that these findings need to be confirmed in additional populations and tested in real-world deployment before such tools could be recommended for widespread clinical use.
Garcia B, Dekker L, Bekker R, Wermer M, van Etten E, van Zwet E, et al.. (2026). Machine learning versus conventional methods for prehospital detection of stroke due to large vessel occlusion or intracranial haemorrhage.. European stroke journal. https://doi.org/10.1093/esj/aakag112