Cardiovascular

ECG-based detection of occlusion myocardial infarction: a dedicated deep neural network versus multimodal large language models and physicians - a retrospective diagnostic accuracy study.

TL;DR

In ECG-based occlusion myocardial infarction detection, the dedicated deep neural network Queen of Hearts showed the highest discrimination (AUC 0.96) and sensitivity (95.8%), outperforming general-purpose large language models and emergency physicians, while general-purpose LLMs showed clinically important limitations including inconsistent outputs and poor calibration.

Key Findings

The dedicated OMI-detection deep neural network Queen of Hearts (QoH) achieved the highest diagnostic discrimination among all interpreters tested.

  • QoH achieved an AUC of 0.96 for OMI detection.
  • ChatGPT followed with an AUC of 0.81, and Gemini performed least well with an AUC of 0.62.
  • The study used 36 twelve-lead ECGs, including 24 with angiographically confirmed acute coronary occlusion and 12 without.
  • QoH provided deterministic output and was queried once, unlike LLMs which were queried five times each.

QoH achieved the highest sensitivity for OMI detection, missing only one of 24 confirmed occlusions.

  • QoH sensitivity was 95.8%, missing one of 24 occlusions.
  • QoH overall accuracy was 86.1%.
  • QoH sensitivity was significantly higher than that of emergency specialists (p = 0.016).
  • Among physicians, specialists performed best with sensitivity of 66.7% and accuracy of 75.0%.

Emergency medicine specialists outperformed residents and LLMs in overall physician-level diagnostic performance for OMI.

  • Specialists achieved accuracy of 75.0%, sensitivity of 66.7%, and specificity of 91.7%.
  • Five emergency medicine specialists and five emergency medicine residents each interpreted all 36 ECGs.
  • All interpreters worked under a standardized moderate-risk acute coronary syndrome clinical scenario.

ChatGPT demonstrated high specificity but low sensitivity for OMI detection, limiting its clinical utility for ruling in occlusion.

  • ChatGPT specificity was 91.7%, but sensitivity was only 54.2%.
  • ChatGPT achieved an AUC of 0.81.
  • Each LLM evaluated the full ECG set five times in separate sessions to assess consistency.

Gemini 3 Pro performed least well overall among all interpreters and showed marked overconfidence in its predictions.

  • Gemini achieved an overall accuracy of 55.6% and an AUC of 0.62.
  • Gemini's Brier score was 0.40, indicating marked overconfidence.
  • The Brier score is a measure of probabilistic calibration, where higher scores indicate worse calibration.

Both LLMs produced variable and inconsistent interpretations when the same ECGs were submitted across repeated sessions.

  • Fleiss' kappa for inter-session reliability ranged from 0.24 to 0.49 across the two LLMs.
  • These kappa values reflect fair to moderate agreement, indicating meaningful inconsistency in LLM outputs.
  • Each LLM evaluated the ECG set five times in separate sessions to quantify this variability.
  • QoH, by contrast, provides deterministic output and was queried only once.

The study design used angiographically confirmed acute coronary occlusion as the reference standard for OMI diagnosis.

  • 36 twelve-lead ECGs were sourced from patients referred for emergent coronary angiography.
  • 24 ECGs were from patients with angiographically confirmed acute coronary occlusion and 12 were from patients without occlusion.
  • This reflects the broader OMI paradigm rather than the traditional STEMI criterion.
  • AUCs were compared using the DeLong test, binary metrics using the exact McNemar test, and reliability using Fleiss' kappa.

The authors concluded that general-purpose LLMs do not currently support autonomous OMI diagnosis and suggest at most an adjunctive, human-supervised role.

  • Clinically important limitations identified in LLMs included inconsistent outputs and poor confidence calibration.
  • The findings support further evaluation of task-specific models for ECG-based occlusion detection.
  • The study is retrospective, and the ECG sample size of 36 is noted as a study limitation implicitly by the design.

What This Means

This research suggests that a specialized artificial intelligence tool called Queen of Hearts (QoH), designed specifically to detect blocked coronary arteries from ECGs, substantially outperforms both general-purpose AI chatbots and experienced emergency doctors at identifying a dangerous heart condition called occlusion myocardial infarction (OMI). OMI occurs when a coronary artery is completely blocked, and rapid diagnosis is critical. The study tested 36 ECGs from real patients—24 with confirmed artery blockages—by having emergency medicine specialists, residents, two leading AI chatbots (ChatGPT and Gemini), and QoH each interpret the tracings. QoH correctly identified 95.8% of true blockages and missed only one, while emergency specialists caught 66.7% and ChatGPT only 54.2%. Gemini performed worst, with accuracy barely above chance and a tendency to express false confidence in its wrong answers. The AI chatbots also gave inconsistent answers when shown the same ECGs on different occasions, which is a serious concern for a diagnostic tool. This inconsistency, combined with poor accuracy and miscalibrated confidence (particularly in Gemini), means these general-purpose chatbots are not reliable enough for independent cardiac diagnosis. By contrast, QoH's output is deterministic—it gives the same answer every time for the same input—and its accuracy was meaningfully better than all human interpreters tested. This research suggests that purpose-built, task-specific AI models like QoH may have genuine clinical value as decision-support tools for time-sensitive ECG interpretation in emergency settings, and warrant further study with larger patient populations. At the same time, it highlights that popular general-purpose AI chatbots, despite their impressive language abilities, are not ready for autonomous use in high-stakes medical decisions like OMI diagnosis and should only be considered, at most, as supplementary tools under close physician supervision.

Have a question about this study?

Citation

Akar E, Kokulu K, Sert E. (2026). ECG-based detection of occlusion myocardial infarction: a dedicated deep neural network versus multimodal large language models and physicians - a retrospective diagnostic accuracy study.. BMC emergency medicine. https://doi.org/10.1186/s12873-026-01716-3