Deep learning transformer-based models using real-world handheld mobile ECG data can predict short-term atrial fibrillation occurrence from normal sinus rhythm recordings with AUROCs of 0.787–0.793, supporting the feasibility of remote AF risk prediction in real-world populations.
Key Findings
Results
Limb-lead transformer models achieved AUROCs of 0.793, 0.785, and 0.787 for 7-, 14-, and 31-day AF occurrence predictions on the internal cohort.
Models were trained on 97,447 labeled mECGs and evaluated on an internal real-world cohort derived from 386,519 mECGs collected from 8,206 users between March 2023 and November 2024.
AF incidences within the 7-, 14-, and 31-day time windows were 18,949, 25,206, and 33,524 respectively.
User-level AUROC for the 31-day prediction was 0.702.
Limb-lead models significantly outperformed lead I models (P<.001), supporting the value of multilead configurations.
Results
Multistage pretraining combining large-scale 12-lead ECGs and proprietary mECGs was essential for model performance, with single-source pretraining yielding substantially lower AUROCs.
The multistage approach used pretraining with 787,257 12-lead ECGs and 202,689 mECGs, followed by fine-tuning on labeled mECGs.
Single-source pretraining with mECGs only yielded an AUROC of 0.555 for 31-day prediction.
Single-source pretraining with 12-lead ECGs only yielded an AUROC of 0.761 for 31-day prediction.
The combined multistage approach achieved an AUROC of 0.787 for 31-day prediction, representing substantial improvements over either single-source approach.
Results
Model performance showed significant disparities by sex and QRS duration but was consistent across age, PR interval, and corrected QT interval subgroups.
AUROC was 0.713 in females versus 0.794 in males (P<.001).
AUROC was 0.583 for QRS duration ≥120 ms versus 0.796 for QRS duration <120 ms (P<.001).
Performance was consistent across age, PR interval, and corrected QT interval subgroups.
These disparities suggest potential limitations in model generalizability across certain patient subpopulations.
Results
In an external cohort proof-of-concept analysis, the 31-day model successfully stratified new-onset AF risk with significantly different survival functions between predicted positive and negative groups.
The external cohort comprised 144 participants.
The model correctly stratified all 5 new-onset AF events in the external cohort.
Survival functions between positively and negatively predicted groups were significantly different (P=.03).
Cox proportional hazards regression yielded a hazard ratio of 1.49 (95% CI 1.06–2.09) per 0.1 increase in model output.
Methods
Handheld mobile ECG devices capable of capturing 6 limb leads were used to collect real-world data from outpatient users for AF prediction model development.
Data were collected from commercially available handheld mECG devices capable of capturing 6 limb leads.
A total of 386,519 mECGs were acquired from 8,206 users between March 2023 and November 2024.
AF occurrence was defined as an AF event within a predefined time window of 7, 14, or 31 days from the date of the NSR recording.
Models were developed for both limb-lead and lead I input configurations using a self-supervised pretraining and domain adaptation approach.
Discussion
The study authors propose that model output may serve as a risk indicator to support opportunistic AF screening in real-world populations.
The authors highlight the potential for remote AF management using single normal sinus rhythm recordings from mobile devices.
AF is described as a common arrhythmia associated with increased risk of stroke and heart failure.
The model is positioned to prompt further clinical evaluation and inform decisions about more intensive monitoring rather than as a standalone diagnostic tool.
The study represents one of the first evaluations of deep learning AF prediction models using mECG in outpatient, real-world settings.
What This Means
This research suggests that artificial intelligence can be used to predict whether someone is at risk of developing atrial fibrillation (an irregular heart rhythm) within the next week or month, using recordings from a handheld mobile heart monitor. The AI models were trained on nearly 800,000 standard hospital ECG recordings combined with over 200,000 recordings from consumer mobile devices, and then tested on data from more than 8,000 real-world users. The models were able to identify at-risk individuals from a single recording taken during a normal heart rhythm, achieving accuracy levels (measured by AUROC) of around 0.79, meaning they correctly distinguished at-risk from low-risk individuals roughly 79% of the time.
The study found that using recordings from multiple leads (six limb leads) rather than just one lead significantly improved prediction accuracy, and that combining different types of training data was critical — using only mobile ECG data or only hospital ECG data alone produced much worse results. However, the models performed less well in women compared to men, and in people with wider-than-normal QRS complexes on their ECG, suggesting the tool may not work equally well for all patients. In a smaller external group of 144 people, the model successfully identified all five individuals who went on to develop new atrial fibrillation within 31 days.
This research matters because atrial fibrillation often goes undetected until it causes a serious complication like stroke, and most current screening methods require clinic visits or specialized equipment. This study suggests that consumer-grade handheld heart monitors combined with AI could enable opportunistic screening in everyday settings, potentially catching high-risk individuals earlier. The authors caution that the model output should be used as a risk indicator to prompt further evaluation rather than as a definitive diagnosis, and that disparities in performance across sexes and certain heart conditions need to be addressed before broad clinical deployment.
Park M, Ahn H, Na Y, Joo S, Lee Y, Han S, et al.. (2026). Prediction of Atrial Fibrillation Occurrence With Handheld Mobile Electrocardiogram: Deep Learning Model Development Using Real-World Data.. JMIR medical informatics. https://doi.org/10.2196/87142