Frontier LLMs are highly reproducible internally but diverge from one another asymmetrically, with ChatGPT and Gemini behaving near-identically while Claude diverges conservatively in unruptured intracranial aneurysm treatment recommendations.
Key Findings
Results
Within-model reproducibility for all three LLMs was almost perfect across five repeated runs of the same clinical vignettes.
Fleiss' κ ranged from 0.837 to 0.860 across Claude, ChatGPT, and Gemini
Each model was run five times per case to assess reproducibility
67 UIA cases were used as clinical vignettes
The almost-perfect internal consistency indicates individual models reliably produce the same recommendation when re-queried
Results
ChatGPT and Gemini showed near-identical pairwise inter-model agreement, while Claude diverged substantially from both.
ChatGPT-Gemini pairwise agreement: Cohen's κ = 0.850 (95% CI 0.71–0.97)
Claude vs. ChatGPT: κ = 0.688
Claude vs. Gemini: κ = 0.667
Inter-model agreement was described as 'asymmetric,' with Claude as the primary source of divergence
Results
The three LLMs produced non-unanimous recommendations in 19.4% of cases, with Claude as the sole outlier in the majority of those disagreements.
Non-unanimous recommendations occurred in 13 of 67 cases (19.4%)
Claude was the sole outlier in 8 of those 13 cases (approximately 61.5%)
In 7 of Claude's 8 outlier cases, Claude was more conservative (i.e., recommended observation rather than intervention)
The other models (ChatGPT and Gemini) were in agreement in these cases
Results
Gemini was significantly more pro-treatment than Claude in head-to-head comparison.
McNemar's test p = 0.0117 for Gemini vs. Claude comparison
Claude was consistently more conservative in its treatment recommendations
This directional difference persisted across the 67 case vignettes
Results
ChatGPT and Gemini demonstrated significant pro-treatment propensity compared to the neurovascular multidisciplinary team (MDT) consensus.
Gemini vs. MDT: McNemar p = 0.0022
ChatGPT vs. MDT: McNemar p = 0.0153
Both models recommended treatment more frequently than the human MDT consensus
MDT consensus served as the clinical ground truth anchor
Results
Claude was significantly more conservative than the Unruptured Intracranial Aneurysm Treatment Score (UIATS) reference standard.
McNemar p = 0.0162 for Claude vs. UIATS
Claude recommended observation (conservative management) more often than UIATS indicated
UIATS was used as an additional validated clinical benchmark alongside MDT consensus
Neither ChatGPT nor Gemini showed a statistically significant deviation from UIATS in the same direction
Methods
The study retrospectively analysed 67 UIA cases referred to a neurovascular service over a one-year period using de-identified clinical vignettes submitted to three frontier LLMs.
Cases were from January–December 2025 referrals to the neurovascular service
Models tested were Claude Opus 4.6, ChatGPT-5.4, and Gemini 3 Pro Thinking
Each vignette was submitted five times to each model to enable within-model reproducibility assessment
Statistical methods included Fleiss' κ (within-model), Cohen's κ and McNemar's test (inter-model and vs. benchmarks)
What This Means
This research suggests that when patients or clinicians ask different AI chatbots — specifically ChatGPT, Gemini, and Claude — whether to treat an unruptured brain aneurysm (a blood vessel bulge in the brain that hasn't yet burst), they may get meaningfully different answers depending on which AI they consult. The study found that each individual AI is very consistent with itself, giving nearly the same answer if asked the same question multiple times. However, ChatGPT and Gemini tend to agree with each other and lean toward recommending treatment, while Claude more often recommends a 'watch and wait' approach. In about 1 in 5 cases, the three AIs did not agree on a recommendation.
When compared to actual decisions made by specialist medical teams and a validated clinical scoring tool, ChatGPT and Gemini recommended treatment more often than the human specialists did, while Claude recommended treatment less often than the scoring tool suggested was appropriate. Neither pattern is necessarily 'right' or 'wrong' in every case, but the disagreements highlight that the AI a patient happens to use could shape their expectations and concerns before they even see a doctor.
This research suggests that as more patients turn to AI tools for health information before seeing specialists, clinicians need to be aware that different AI systems may have given a patient conflicting or biased information. The authors call for further research into why these differences exist across AI models and how patients emotionally respond to AI-generated medical recommendations.
Narayan A, Mariajoseph F, Farooq M, Praeger A, Chandra R, Sher I, et al.. (2026). Inter-Model Variability of ChatGPT, Gemini and Claude in Treatment Recommendations for Unruptured Intracranial Aneurysms.. Neurosurgical review. https://doi.org/10.1007/s10143-026-04498-1