When enough items are included, which specific items make up the frailty index has negligible influence on prediction, and combining items into a single score discards some prognostic information but maintains the generalisability that makes the FI practical.
Key Findings
Results
Models using all 46 individual frailty deficit items as separate predictors outperformed composite frailty index approaches for both all-cause and CVD-specific mortality.
C-index for 46 individual items was 0.71 (95% CI: 0.69-0.74) compared to 0.64-0.67 for composite FI approaches.
This was observed for both all-cause and CVD-specific mortality outcomes.
Random survival forest models were used alongside Cox models for these comparisons.
The sample included 3,669 adults living with cardiovascular disease from NHANES (1999-2018).
Results
Importance-ranked frailty index performance peaked at approximately 10 items and then declined as more items were added.
Items were ranked by permutation importance within random survival forest models.
At each item count k (ranging from 1 to 46), the top k items were combined into a single FI score.
Performance decline after ~10 items suggests that adding lower-ranked items dilutes rather than improves the composite score.
The authors interpret this as short importance-ranked FIs benefiting from 'avoiding dilution rather than capturing an optimal item set.'
Results
Randomly built frailty indices improved steadily with more items and converged with importance-ranked composites by approximately 35 items for all-cause mortality and 20 items for CVD-specific mortality.
At each item count k, 200 FIs were each built from randomly selected k items.
Convergence between random and importance-ranked FIs occurred at ~35 items for all-cause mortality.
Convergence occurred at ~20 items for CVD-specific mortality.
This convergence indicates that item selection becomes negligible when a sufficient number of items are included.
Results
When enough items are included, which specific items make up the frailty index has negligible influence on mortality prediction.
This finding supports the frailty index assumption that deficits are interchangeable when item counts are sufficient.
The specific threshold differed by outcome: ~35 items for all-cause mortality and ~20 items for CVD-specific mortality.
Analyses using imputation produced consistent results, supporting robustness of findings.
The study used a 46-item FI derived from NHANES data spanning 1999-2018.
Results
Combining frailty deficits into a single composite score discards some prognostic information compared to using items individually, but maintains generalisability.
Individual items as separate predictors achieved C-index of 0.71 (0.69-0.74), while composite FI scores reached only 0.64-0.67.
The authors describe this as a 'trade-off' that 'maintains the generalisability that makes the FI practical.'
The finding suggests that machine learning approaches using individual items can extract more predictive signal than composite scoring.
Despite the performance gap, the composite FI approach retains clinical utility through its simplicity and generalisability.
Discussion
Short importance-ranked frailty indices are outcome-specific and sample-dependent, limiting their generalisability.
Item importance rankings differed between all-cause and CVD-specific mortality outcomes.
The authors conclude that 'any such selection is outcome-specific and sample-dependent.'
This finding challenges the motivation for using machine learning to create shorter FIs for clinical settings.
The study was conducted specifically in adults with cardiovascular disease from NHANES, which may limit generalisability to other populations.
Methods
The study was conducted in 3,669 adults living with cardiovascular disease using a 46-item frailty index from NHANES data collected between 1999 and 2018.
Data source was the National Health and Nutrition Examination Survey (NHANES), cycles 1999-2018.
All participants had cardiovascular disease.
The frailty index contained 46 items.
Both Cox regression and random survival forest models were employed.
Outcomes included all-cause mortality and CVD-specific mortality.
What This Means
This research suggests that when it comes to measuring frailty using a frailty index (FI) — a tool that counts health deficits like chronic conditions, disabilities, and symptoms — the specific items chosen matter less than how many items are included. The study analyzed nearly 3,700 adults with cardiovascular disease and found that once around 20-35 health deficit items were included in the FI, randomly choosing which items to include performed just as well as carefully selecting the 'best' items for predicting death. This supports a core assumption of the frailty index: that individual health problems are interchangeable, and what matters is the total burden of deficits rather than which specific ones are present.
However, the study also found that using all 46 health deficit items as separate predictors (rather than combining them into a single score) produced better predictions of mortality than any composite frailty score. This means that combining items into one number — while convenient and broadly useful — does lose some predictive information. Additionally, shorter frailty indices built by selecting only the most 'important' items peaked in performance at around 10 items and then got worse as more items were added, suggesting their apparent advantage comes from avoiding less informative items rather than from identifying a truly optimal set of deficits. Importantly, which items ranked as most important differed depending on whether all-cause or CVD-specific death was the outcome being predicted.
This research suggests that efforts to create shorter, machine-learning-optimized frailty indices for clinical use may have limited advantages over standard comprehensive frailty assessment, especially because the 'best' short item set will vary depending on what health outcome is being predicted and in which patient population. The traditional approach of collecting a broad range of health deficits and combining them into a single frailty score remains practical and generalizable, even if it sacrifices some predictive precision compared to more complex approaches.
Quach J, Theou O, Rockwood K, Kehler S, Blodgett J. (2026). Frailty index deficit interchangeability: an empirical test using random survival forests in adults with cardiovascular disease.. Age and ageing. https://doi.org/10.1093/ageing/afag262