
AI Model Outperforms Apnea-Hypopnea Index in Predicting Sleep Apnea Mortality Risk
Key Takeaways
- Training on raw EEG, ECG, airflow, oximetry, and related channels generated high-dimensional embeddings that clustered into five reproducible risk groups across 9608 polysomnograms.
- Fully adjusted models showed the highest-risk group had markedly elevated all-cause mortality (HR 2.38) and increased MACE, heart failure, MI, atrial fibrillation, cognitive impairment, and epilepsy.
A foundation model built on routine sleep studies outperformed the apnea-hypopnea index (AHI) in predicting mortality and cardiovascular risk in patients with sleep apnea.
A foundation model trained on routine
The apnea-hypopnea index (AHI)—a measure of breathing interruptions per hour of sleep—has anchored the diagnosis and severity grading of obstructive sleep apnea (OSA) for decades. But AHI reduces an entire night of multichannel physiologic recording to a single number, and a growing body of evidence suggests that number carries limited prognostic value.
In the new study, AHI severity categories showed no significant association with all-cause mortality in the study cohort, while an
Model Built From Routine Polysomnograms
Researchers from Cleveland Clinic, IBM Research, and Yale School of Medicine trained a transformer-based foundation model on 9608 polysomnograms (PSGs) from 9297 patients in Cleveland Clinic's Sleep Signals, Testing, and Reports Linked to Patient Traits (STARLIT-10K) registry, linked to electronic medical records spanning a mean observation period of 14.5 years. Unlike conventional approaches that rely on technologist-scored summary metrics, the model ingested raw signals across electroencephalogram (EEG), electrooculogram, electromyogram, electrocardiogram (ECG), airflow, oximetry, and other channels, generating high-dimensional physiologic embeddings for each patient.
Clustering those embeddings produced 5 stable risk groups (RG1 through RG5). The findings were not incidental to respiratory signals alone: sensitivity analyses that zeroed out EEG or ECG inputs substantially altered cluster assignments, indicating the model draws on neurologic and cardiac information that AHI does not capture at all.
Risk Groups Reveal Outcomes AHI Missed
The gradient in outcomes across groups was pronounced. In fully adjusted models, the highest-risk group (RG5) carried an HR of 2.38 (P = 1.20×10⁻⁷) for all-cause mortality relative to the lowest-risk group, along with elevated risk for major adverse cardiovascular events (HR, 1.64), heart failure (HR, 1.65), myocardial infarction (HR, 1.84), atrial fibrillation (HR, 2.23), cognitive impairment (HR, 1.93), and epilepsy (HR, 2.40). These associations held after adjusting for AHI itself, meaning the risk captured by the model was not simply a proxy for apnea severity.
A Sankey diagram in the study illustrated why this matters clinically: patients with severe AHI were scattered across multiple risk groups, and the highest-risk cluster included patients spanning the full range of AHI severities. In other words, AHI alone would misclassify a meaningful share of high-risk patients as low risk, and vice versa.
The model's findings were not confined to a single institution. Applying the same pipeline to the independent Sleep Heart Health Study cohort reproduced the risk gradient for mortality and incident heart failure, despite that cohort's lower-resolution recordings.
Findings Reflect a Broader Diagnostic Shift
The findings add to a growing body of research repositioning sleep diagnostics around physiologically grounded, machine-derived metrics rather than conventional frequency-based indices. A deep learning model
Together, these developments suggest sleep medicine is moving toward automatically derived measures of arousal burden, hypoxic burden, and now whole-night physiologic risk that may better capture disease severity than AHI alone.
Does the Model Close Sex-Based Diagnostic Gaps?
One notable secondary finding concerns sex-based diagnostic
The foundation model, by contrast, demonstrated consistent prognostic value across both cohorts without the demographic restrictions that limited AHI's utility, suggesting AI-based stratification could help close a longstanding blind spot in how OSA risk is assessed in women.
Study Limitations and Next Steps
This is a discovery and validation-phase study, not a clinically deployed diagnostic tool; the risk groups were identified retrospectively and have not yet been tested as part of an actual care pathway. The authors caution that the retrospective design limits causal inference and that objective data on continuous positive airway pressure adherence were unavailable. Sensitivity analyses excluding patients with a positive airway pressure prescription produced consistent results, but incomplete treatment data could still bias associations toward the null.
In addition, the Cleveland Clinic dataset itself is not publicly available, and the proprietary production code was not released. However, the authors state the methodology is documented in sufficient detail for independent reconstruction. Still, prospective validation will be needed before this kind of stratification could inform clinical workflows, referral patterns, or trial enrollment.
The findings open new directions for future research inquiries, according to the authors.
“Patient-reported outcomes and symptom measures were not included here but remain central to sleep medicine,” they wrote. “Future work integrating subjective data with embeddings may yield more informative phenotypes.”
References
- Bilal E, Araujo MLD, Beck KL, et al. A foundation model for sleep-based risk stratification and clinical outcomes. Nat Commun. 2026;17:7603. doi:10.1038/s41467-026-75326-9
- Grossi G. Deep learning model matches expert agreement in sleep arousal scoring. AJMC®. July 17, 2026. Accessed August 13, 2026.
https://www.ajmc.com/view/deep-learning-model-matches-expert-agreement-in-sleep-arousal-scoring




