DLIR was feasible for CACS measurement and showed high overall concordance with FBP and IR, but near-perfect correlation should not be interpreted as interchangeability, because DLIR showed systematic underestimation and non-negligible individual-level disagreement from FBP, particularly in patients with low calcium burden or scores near CAC risk-category thresholds.
Key Findings
Results
In the phantom study, calcium scores were comparable across all three reconstruction methods (FBP, IR, and DLIR).
A cardiac phantom was used to obtain calcium scores for FBP, IR (adaptive statistical iterative reconstruction-V with 80% blending), and DLIR (GE TrueFidelity at high strength).
No significant differences in calcium scores were observed between reconstruction methods in the phantom setting.
The Agatston method was applied to all three techniques to obtain coronary artery calcium scores.
Results
In the clinical study, DLIR produced significantly lower calcium scores than FBP, while the difference between IR and DLIR was not statistically significant.
113 patients who underwent coronary computed tomography angiography were retrospectively evaluated.
DLIR had significantly lower scores than FBP (p < 0.001).
The comparison between IR and DLIR was statistically non-significant (p = 0.064).
This suggests DLIR systematically underestimates calcium scores relative to FBP.
Results
Weighted kappa coefficients indicated high agreement in total CACS classification across reconstruction methods.
High agreement in total CACS classification was observed across FBP, IR, and DLIR using weighted kappa analysis.
Among 113 patients, 109 were classified into the same CAC-DRS category by both FBP and DLIR.
DLIR underclassified four patients compared to FBP, yielding a false-negative rate of 4.65%.
Results
Bland-Altman analysis demonstrated non-negligible individual-level disagreement between FBP and DLIR.
The mean difference between FBP and DLIR was 14.5.
The 95% limits of agreement ranged from -52.1 to 81.1.
This level of individual disagreement was characterized as 'non-negligible.'
The disagreement was considered potentially clinically relevant near established CAC risk-category thresholds, particularly at 0, 100, and 300.
Results
DLIR showed a false-negative rate of 4.65% for CAC-DRS category classification compared to FBP.
Four out of 113 patients were underclassified by DLIR relative to FBP.
This corresponds to a false-negative rate of 4.65% (4/86 patients with non-zero CAC by FBP, or 4/113 total).
Underclassification risk was noted to be particularly relevant for patients with low calcium burden or scores near risk-category thresholds.
Conclusions
Despite high overall concordance, FBP and DLIR should not be considered interchangeable for coronary calcium scoring.
The authors explicitly state that 'near-perfect correlation should not be interpreted as interchangeability.'
DLIR showed systematic underestimation relative to FBP.
Individual-level disagreement was particularly concerning near CAC risk-category thresholds of 0, 100, and 300.
The study recommends caution in clinical settings, especially for patients near threshold values where risk category reclassification could occur.
What This Means
This research evaluated whether a newer, AI-based method for reconstructing CT scan images — called deep learning-based image reconstruction (DLIR) — could reliably measure calcium deposits in coronary arteries, compared to two older standard methods (filtered back projection, or FBP, and iterative reconstruction, or IR). Calcium scoring from CT scans is an important tool for assessing heart disease risk, and the researchers wanted to know if DLIR could be used interchangeably with established techniques. They tested these methods both in a physical phantom (a device that mimics the body) and in data from 113 real patients.
The study found that while DLIR performed similarly to the older methods in the phantom test and showed high overall agreement with FBP and IR in patients, it consistently produced lower calcium scores than FBP in real patients. When looking at individual patients, the differences between FBP and DLIR were sometimes large enough to matter clinically — meaning a patient could be placed into a different risk category depending on which reconstruction method was used. Specifically, 4 out of 113 patients were assigned to a lower risk category by DLIR than by FBP, representing a false-negative rate of about 4.65%. This risk of misclassification was greatest for patients whose calcium scores fell near the key thresholds used to guide clinical decisions (scores of 0, 100, and 300).
This research suggests that while DLIR is technically feasible for coronary calcium scoring and agrees well with traditional methods at the group level, the two approaches cannot simply be swapped for one another without caution. Patients with borderline calcium scores — those near clinically important thresholds — may be at risk of being assigned to a lower-risk category if DLIR is used instead of the traditional FBP method. Clinicians and imaging centers considering a switch to DLIR should be aware of this systematic underestimation and its potential impact on individual patient risk assessment.
Kim J, Kim T, Lee S, Nam J, Park C. (2026). Feasibility of deep learning-based image reconstruction using TrueFidelity for coronary calcium scoring.. PloS one. https://doi.org/10.1371/journal.pone.0358342