A machine-learning model commissioned to identify undiagnosed hypertension in North West London showed overall sensitivity of 62.7% and specificity of 60.7%, with considerable demographic variation in performance and low positive predictive value, performing comparably to a more interpretable logistic regression model.
Key Findings
Results
The NWL hypertension machine-learning model demonstrated moderate overall predictive performance with sensitivity of 62.7% and specificity of 60.7%.
Overall sensitivity was 62.7% (95% CI 62.5-62.8) and specificity was 60.7% (95% CI 60.5-60.8).
The study cohort comprised 1,802,920 individuals aged 16 years or older registered with a general practice in North West London with no prior hypertension diagnosis, from May 2023 to May 2024.
Positive predictive value (PPV) ranged from 31.5% (95% CI 31.2-31.8) to 42.9% (95% CI 42.5-43.2).
Negative predictive value (NPV) ranged from 77.6% (95% CI 77.2-77.9) to 84.9% (95% CI 84.7-85.2).
Model predictions were assessed against recorded hypertension status using both medical diagnoses and blood pressure records.
Results
Model sensitivity was higher in older adults and Black patients, while specificity was higher in younger adults, female patients, and White patients.
Logistic regression models were used to assess sensitivity and specificity variation by sociodemographic groups.
Sensitivity was higher among Black patients compared to other ethnic groups.
Specificity was higher in female patients and White patients.
Socioeconomic deprivation was associated with higher sensitivity but lower specificity, with effects plateauing in the two least deprived quintiles.
Overall sensitivity was higher for those living in areas of higher socioeconomic deprivation, while specificity was lower.
The effects plateaued in the 2 least deprived quintiles of deprivation, which were comparable in both sensitivity and specificity.
This pattern suggests the model may perform more consistently in less deprived populations.
Results
Model predictions varied substantially by age, with nearly all individuals aged 70-79 predicted to have hypertension and very few younger adults predicted positive.
96.2% (58,951/61,281) of those aged 70 to 79 were predicted to have hypertension.
Only 0.08% of those aged 20 to 39 were predicted to have the condition.
This extreme variation suggests the model's predictions are heavily age-driven.
Results
The machine-learning model's performance was comparable to a more interpretable logistic regression model.
The study directly compared the NWL ML model against a more interpretable regression approach.
No meaningful difference in predictive performance was found between the two approaches.
The authors recommend that 'where there is no loss in performance, more parsimonious, transparent models be selected for prediction in health care settings.'
The comparability of performance raises questions about the added value of complex ML models over simpler alternatives in this context.
Results
Despite relatively good performance for identifying those without hypertension, the model left a significant proportion of true hypertension cases undetected.
The low positive predictive value (31.5% to 42.9%) means a majority of those flagged by the model did not actually have undiagnosed hypertension.
A 'significant proportion of true cases remained undetected' according to the authors.
The authors note this limits 'opportunities for early intervention' in the context of hypertension as 'a leading preventable cause of cardiovascular disease.'
What This Means
This research evaluated a computer model built by the UK's National Health Service (NHS) to identify people who likely have high blood pressure (hypertension) but have not yet been diagnosed. The model was tested on nearly 1.8 million adults registered with general practices in North West London. The researchers found that the model correctly identified about 63% of people who actually had undiagnosed hypertension (sensitivity) and correctly identified about 61% of people who did not have the condition (specificity). Importantly, when the model flagged someone as likely having hypertension, that prediction was correct only about 32-43% of the time — meaning most people the model flagged would turn out not to have the condition after testing.
The model did not perform equally across different groups of people. It was better at detecting hypertension in older adults and Black patients, but more likely to incorrectly flag people in these groups as well. It was more precise (fewer false alarms) in younger adults, women, and White patients. People living in more deprived (poorer) areas were more likely to be detected if they had hypertension, but also more likely to be incorrectly flagged if they did not. A striking finding was that the model predicted nearly all people aged 70-79 had hypertension, while almost no one aged 20-39 was flagged — suggesting age plays a dominant role in the model's predictions. Notably, this complex machine-learning model performed no better than a much simpler, more transparent statistical model.
This research suggests that while such predictive models can play a role in identifying people at risk of undiagnosed hypertension, their practical usefulness is limited by a high rate of false positives and unequal performance across demographic groups. The findings point to the need for tailored screening approaches for different population groups to ensure fairness, and support the use of simpler, more understandable models when they perform just as well. The results can help health services decide where and how to best direct hypertension screening efforts and whether the resources needed to act on model predictions are justified.
Ihenetu G, Alkhatib A, Novov V, Beaney T, Majeed A, Aylin P, et al.. (2026). Evaluation of a National Health Service Machine-Learning Model for Hypertension Case-Finding: Retrospective Cohort Study.. Journal of medical Internet research. https://doi.org/10.2196/87084