News|Articles|August 13, 2026

AI Model Outperforms Apnea-Hypopnea Index in Predicting Sleep Apnea Mortality Risk

Fact checked by: Brooke McCormick
Listen
0:00 / 0:00

Key Takeaways

  • Training on raw EEG, ECG, airflow, oximetry, and related channels generated high-dimensional embeddings that clustered into five reproducible risk groups across 9608 polysomnograms.
  • Fully adjusted models showed the highest-risk group had markedly elevated all-cause mortality (HR 2.38) and increased MACE, heart failure, MI, atrial fibrillation, cognitive impairment, and epilepsy.
SHOW MORE

A foundation model built on routine sleep studies outperformed the apnea-hypopnea index (AHI) in predicting mortality and cardiovascular risk in patients with sleep apnea.

A foundation model trained on routine sleep study data identified patient subgroups with markedly different long-term health risks that the standard clinical measure for sleep apnea failed to detect, according to a study published in Nature Communications.1

The apnea-hypopnea index (AHI)—a measure of breathing interruptions per hour of sleep—has anchored the diagnosis and severity grading of obstructive sleep apnea (OSA) for decades. But AHI reduces an entire night of multichannel physiologic recording to a single number, and a growing body of evidence suggests that number carries limited prognostic value.

In the new study, AHI severity categories showed no significant association with all-cause mortality in the study cohort, while an artificial intelligence (AI) model built on the same underlying data identified a patient subgroup with more than double the 5-year mortality risk of the lowest-risk group.

Model Built From Routine Polysomnograms

Researchers from Cleveland Clinic, IBM Research, and Yale School of Medicine trained a transformer-based foundation model on 9608 polysomnograms (PSGs) from 9297 patients in Cleveland Clinic's Sleep Signals, Testing, and Reports Linked to Patient Traits (STARLIT-10K) registry, linked to electronic medical records spanning a mean observation period of 14.5 years. Unlike conventional approaches that rely on technologist-scored summary metrics, the model ingested raw signals across electroencephalogram (EEG), electrooculogram, electromyogram, electrocardiogram (ECG), airflow, oximetry, and other channels, generating high-dimensional physiologic embeddings for each patient.

Clustering those embeddings produced 5 stable risk groups (RG1 through RG5). The findings were not incidental to respiratory signals alone: sensitivity analyses that zeroed out EEG or ECG inputs substantially altered cluster assignments, indicating the model draws on neurologic and cardiac information that AHI does not capture at all.

Risk Groups Reveal Outcomes AHI Missed

The gradient in outcomes across groups was pronounced. In fully adjusted models, the highest-risk group (RG5) carried an HR of 2.38 (P = 1.20×10⁻⁷) for all-cause mortality relative to the lowest-risk group, along with elevated risk for major adverse cardiovascular events (HR, 1.64), heart failure (HR, 1.65), myocardial infarction (HR, 1.84), atrial fibrillation (HR, 2.23), cognitive impairment (HR, 1.93), and epilepsy (HR, 2.40). These associations held after adjusting for AHI itself, meaning the risk captured by the model was not simply a proxy for apnea severity.

A Sankey diagram in the study illustrated why this matters clinically: patients with severe AHI were scattered across multiple risk groups, and the highest-risk cluster included patients spanning the full range of AHI severities. In other words, AHI alone would misclassify a meaningful share of high-risk patients as low risk, and vice versa.

The model's findings were not confined to a single institution. Applying the same pipeline to the independent Sleep Heart Health Study cohort reproduced the risk gradient for mortality and incident heart failure, despite that cohort's lower-resolution recordings.

Findings Reflect a Broader Diagnostic Shift

The findings add to a growing body of research repositioning sleep diagnostics around physiologically grounded, machine-derived metrics rather than conventional frequency-based indices. A deep learning model matched or exceeded expert agreement on sleep arousal scoring across 10 independent scorers, another effort aimed at extracting more clinically meaningful signal from PSG data than manual, single-metric scoring allows.2

Together, these developments suggest sleep medicine is moving toward automatically derived measures of arousal burden, hypoxic burden, and now whole-night physiologic risk that may better capture disease severity than AHI alone.

Does the Model Close Sex-Based Diagnostic Gaps?

One notable secondary finding concerns sex-based diagnostic disparities. AHI-based risk assessment has historically performed better in men than in women; in the original Sleep Heart Health Study analyses, AHI-based associations with mortality and incident heart failure reached significance only in men.1

The foundation model, by contrast, demonstrated consistent prognostic value across both cohorts without the demographic restrictions that limited AHI's utility, suggesting AI-based stratification could help close a longstanding blind spot in how OSA risk is assessed in women.

Study Limitations and Next Steps

This is a discovery and validation-phase study, not a clinically deployed diagnostic tool; the risk groups were identified retrospectively and have not yet been tested as part of an actual care pathway. The authors caution that the retrospective design limits causal inference and that objective data on continuous positive airway pressure adherence were unavailable. Sensitivity analyses excluding patients with a positive airway pressure prescription produced consistent results, but incomplete treatment data could still bias associations toward the null.

In addition, the Cleveland Clinic dataset itself is not publicly available, and the proprietary production code was not released. However, the authors state the methodology is documented in sufficient detail for independent reconstruction. Still, prospective validation will be needed before this kind of stratification could inform clinical workflows, referral patterns, or trial enrollment.

The findings open new directions for future research inquiries, according to the authors.

“Patient-reported outcomes and symptom measures were not included here but remain central to sleep medicine,” they wrote. “Future work integrating subjective data with embeddings may yield more informative phenotypes.”

References

  1. Bilal E, Araujo MLD, Beck KL, et al. A foundation model for sleep-based risk stratification and clinical outcomes. Nat Commun. 2026;17:7603. doi:10.1038/s41467-026-75326-9
  2. Grossi G. Deep learning model matches expert agreement in sleep arousal scoring. AJMC®. July 17, 2026. Accessed August 13, 2026. https://www.ajmc.com/view/deep-learning-model-matches-expert-agreement-in-sleep-arousal-scoring