Cardiovascular

Artificial intelligence-integrated multimodal retinal imaging for early detection and risk stratification of systemic vascular and neurodegenerative diseases.

TL;DR

RetinalVNG-Net, a multimodal deep learning framework integrating fundus photography, OCT, and clinical metadata, achieved macro-averaged AUC-ROC of 0.944 on external validation for simultaneous risk stratification of hypertensive retinopathy, diabetic retinopathy, and neurodegenerative-associated retinal changes, significantly outperforming single-modality models.

Key Findings

RetinalVNG-Net achieved strong internal cross-validation performance across four classification categories.

  • Internal five-fold cross-validation yielded macro-averaged AUC-ROC of 0.957 (±0.008)
  • Sensitivity was 0.913, specificity 0.941, and F1 score 0.908 across four classes
  • The development set comprised 1,887 subjects from Centers A and B used for cross-validation
  • A stratified 15% subset (n=333) was reserved separately for hyperparameter tuning

RetinalVNG-Net demonstrated generalizability on a geographically distinct, device-heterogeneous external test set.

  • External test set comprised 520 subjects from Center C, which was geographically distinct and device-heterogeneous
  • Macro-averaged AUC-ROC on the external test set reached 0.944 (95% CI 0.922–0.958)
  • Per-class AUCs were 0.963 for hypertensive retinopathy, 0.951 for diabetic retinopathy, 0.924 for neurodegenerative changes, and 0.938 for controls

The multimodal RetinalVNG-Net significantly outperformed the best single-modality fundus model.

  • Macro-AUC of 0.944 for RetinalVNG-Net versus 0.901 for the best single-modality fundus model
  • The difference was statistically significant (p < 0.001)
  • The single-modality comparison model used fundus photography only

RetinalVNG-Net employs a multimodal architecture integrating three distinct data streams with cross-modal attention fusion.

  • The framework integrates fundus photography (encoded via RETFound ViT-Large), OCT data via a dual-stream branch (ResNet-3D-18 for volumetric B-scans and 2D-CNN for layer thickness maps), and clinical metadata via a tabular transformer
  • Modalities are fused via cross-modal attention
  • An auxiliary regression head outputs a continuous Retinal Biological Age Gap (RBAG) score as an interpretable severity biomarker

The study was a retrospective multi-center design encompassing 2,740 subjects across three independent ophthalmology centers.

  • Total sample size was 2,740 subjects from three centers collected between January 2019 and December 2023
  • Centers A and B (n=2,220) constituted the development set; Center C (n=520) served as the external test set
  • The study targeted simultaneous risk stratification of hypertensive retinopathy, diabetic retinopathy, neurodegenerative-associated retinal changes, and controls

The authors note that prospective longitudinal studies are required to establish the value of RetinalVNG-Net for early or predictive detection.

  • The paper states: 'Prospective longitudinal studies would be required to establish value for early or predictive detection'
  • The current study supports 'further evaluation as a tool for risk stratification of retinal manifestations associated with systemic vascular and neurodegenerative disease'
  • The framework is described as demonstrating 'promising robustness and generalizability' rather than proven clinical utility

What This Means

This research describes the development and testing of an artificial intelligence system called RetinalVNG-Net that analyzes eye scans to help identify signs of three different diseases: high blood pressure-related eye damage, diabetic eye disease, and changes associated with neurodegenerative conditions like Alzheimer's disease. The system combines three types of information — color photographs of the retina, detailed 3D scans of retinal layers (OCT), and patient health data — to classify patients into risk categories. Tested on 2,740 patients from multiple medical centers, the AI achieved strong accuracy, correctly distinguishing between disease types about 94% of the time on patients from a new, previously unseen hospital, and performed significantly better than a system using only one type of eye image. This research suggests that combining multiple types of retinal imaging with patient health data in a single AI framework can improve the ability to detect and categorize eye-related signs of serious systemic diseases more accurately than any single imaging approach alone. The system also produces a 'Retinal Biological Age Gap' score, which could potentially serve as a simple number summarizing how much the retina appears older or more damaged than expected for a patient's age. The fact that the system performed well across different hospitals using different equipment suggests it may be adaptable to real-world clinical settings. However, the authors themselves emphasize important limitations: because this was a retrospective study (looking back at existing records), it cannot yet be concluded that the system can predict disease before symptoms appear or improve patient outcomes. The researchers call for prospective studies — following patients forward in time — to determine whether the tool has genuine value for early or predictive detection of these conditions.

Have a question about this study?

Citation

Du K, Zheng Y, Wang Z. (2026). Artificial intelligence-integrated multimodal retinal imaging for early detection and risk stratification of systemic vascular and neurodegenerative diseases.. Frontiers in neurology. https://doi.org/10.3389/fneur.2026.1885303