SDoH explain 70% to 80% of the cross-tract variation in hypertension and diabetes prevalence in NYC, with the top 10 SDoH centering on socioeconomic disadvantage, built environment, and commute time.
Key Findings
Results
Social determinants of health explain 70% to 80% of cross-tract variation in hypertension prevalence across NYC census tracts.
R2 ranged from 66.2% to 79.5% for hypertension models
Root mean square error ranged from 2.6 to 3.2 for hypertension models
Models were run for both age-adjusted and unadjusted prevalence
Analysis used Extreme Gradient Boosting machine learning approach
Results
Social determinants of health explain 70% to 80% of cross-tract variation in diabetes prevalence across NYC census tracts.
R2 ranged from 72.9% to 79.3% for diabetes models
Root mean square error ranged from 2.1 to 2.2 for diabetes models
Models were run for both age-adjusted and unadjusted prevalence
Results were slightly more consistent across model types for diabetes than hypertension
Results
The top 10 SDoH variables contributed to the majority of each model's predictive power for both hypertension and diabetes.
Top 10 SDoH contributed 73% (age-adjusted) and 70% (unadjusted) of hypertension model prediction
Top 10 SDoH contributed 73% (age-adjusted) and 71% (unadjusted) of diabetes model prediction
Contributions were measured using normalized mean absolute Shapley Additive Explanations (SHAP) values
SHAP values allowed identification of the relative importance of each social risk factor
Results
The top 10 SDoH most strongly associated with hypertension and diabetes prevalence centered on domains of socioeconomic disadvantage, built environment, and commute time.
Domains identified through SHAP values from the Extreme Gradient Boosting models
Findings were consistent across both hypertension and diabetes outcomes
Both age-adjusted and unadjusted models pointed to similar top SDoH domains
These domains were identified across all census tracts citywide, not just high-prevalence areas
Results
Neighborhood-specific SDoH associated with hypertension and diabetes varied across census tracts and boroughs in the highest prevalence quintile.
The study identified neighborhood-specific SDoH for census tracts in the highest prevalence quintile for both conditions
Variation was observed across NYC boroughs, suggesting heterogeneity in local social risk drivers
This finding supports the need for place-based, targeted interventions rather than uniform citywide approaches
The retrospective cohort included clinical data on 3.2 million NYC residents integrated with census tract-level SDoH data
Methods
The study used a retrospective cohort of 3.2 million NYC residents with clinical data integrated with census tract-level SDoH data.
Clinical data were linked to SDoH data at the census tract level
Analysis was conducted at the census tract geographic unit across NYC
The study was designed to support the HealthyNYC initiative launched in 2025
HealthyNYC aims to reduce cardiovascular disease and diabetes deaths by 5% by 2030
What This Means
This research suggests that where people live and the social conditions of their neighborhoods — such as poverty, the built environment, and commute times — are strongly linked to how common high blood pressure (hypertension) and diabetes are across New York City neighborhoods. Using machine learning on health data from 3.2 million NYC residents combined with neighborhood-level social data, the researchers found that these social factors account for roughly 70–80% of the differences in hypertension and diabetes rates between census tracts. Just the top 10 social factors identified were enough to explain about 70–73% of what the models predicted for each condition.
The study also found that the specific social factors most important in high-burden neighborhoods varied depending on which part of the city was examined, suggesting that different neighborhoods may need different targeted approaches rather than a one-size-fits-all citywide strategy. This research was designed to support New York City's HealthyNYC initiative, which aims to cut deaths from cardiovascular disease and diabetes by 5% by 2030.
The practical implication of this research is that addressing social conditions like economic disadvantage, neighborhood infrastructure, and transportation burdens may be important levers for reducing the prevalence of hypertension and diabetes in New York City. The findings provide a data-driven map for city planners and public health officials to prioritize specific neighborhoods and tailor interventions to the unique social risk factors present in each community.
Adamson E, Li H, Xu Z, Tanner D, Ling W, Qi Y, et al.. (2026). Using Machine Learning to Identify Social Risk Factors of Hypertension and Diabetes in New York City: Evidence to Support the HealthyNYC Initiative.. Journal of the American Heart Association. https://doi.org/10.1161/JAHA.125.049029