News|Articles|September 1, 2026

Lung Cancer Risk Models Are Multiplying Faster Than They Can Be Validated

Fact checked by: Maggie L. Shaw

Key Takeaways

  • Of 91 lung cancer risk models developed since 2020, fewer than half underwent external validation, and calibration was inconsistently reported.
  • Discrimination ran from an AUC near 0.70 to above 0.90, but external performance fell below internal, pointing to overfitting and population differences.
SHOW MORE

Review finds 91 new lung cancer screening and pulmonary nodule AI tools, but few get external validation or calibration, risking flawed LDCT decisions.

Ninety-one new risk models have been built to tell clinicians who belongs in a lung cancer screening program and which of the nodules found there are dangerous.1 Fewer than half have ever been tested outside the population that produced them.

Why Screening Eligibility Became a Modeling Problem

Lung cancer remains the deadliest malignancy worldwide, accounting for roughly 2.5 million new cases and 1.8 million deaths in 2022, and it is still most often caught late, when 5-year survival sits below 18%. Detection at earlier stages could lift that figure to 80% or higher, and screening programs have expanded across North America and Europe in recent years.

The US Preventive Services Task Force gave an annual low-dose CT (LDCT) screening a grade B recommendation in 2021 for adults aged 50 to 80 years with a 20 pack-year smoking history who currently smoke or quit within the past 15 years, lowering both the starting age and the pack-year threshold it had set in 2013.2 Fixed eligibility criteria of that kind still cover many people who will never develop lung cancer while excluding a substantial share of those who die from it. Risk-based selection was proposed to close that gap, and Canadian province-wide screening programs began using the PLCOm2012 risk model in 2022.

Two databases were searched for models published between January 2020 and January 2026, an update to a prior review that had identified 75 models published between 1985 and early 2020.1 Of 2462 records retrieved, 91 studies met criteria after full-text review. The review protocol was registered with PROSPERO, and the search strategy was documented before it was run.

Fifty-six of the models were designed for screening selection before a scan, 30 of which incorporated a biomarker such as protein or genetic data, and 35 were built to classify malignancy risk in nodules already detected on LDCT. Development and validation samples spanned North America, Europe, and Asia and drew on clinical trials, observational studies, and health-system registries. Regression-based methods predominated, although machine learning and deep learning approaches appeared with increasing frequency.

Discrimination Held Up Better Than Calibration

Discrimination ranged from moderate, with areas under the curve (AUCs) near 0.70, to excellent, above 0.90, and models enhanced with biomarkers or quantitative imaging features frequently outperformed those relying on clinical variables alone. External discrimination ran generally lower than the corresponding internal values where both were reported, a pattern the authors attributed to overfitting and to differences between development and validation populations.

Calibration, or whether predicted risks match observed outcomes, was assessed far less often. Among the screening-selection models built without biomarkers, calibration was evaluated internally in 7 studies and externally in only 3. Among the 30 models that added a biomarker, external validation was reported for 10 and calibration for 11, most of it graphical rather than quantitative. Among the nodule-classification models, 14 reported internal calibration and 5 reported external calibration. Fewer than half of the 91 models underwent robust external validation.

An exploratory meta-regression of study-level AUC values identified no statistically significant predictors of performance. Among screening-selection models, machine learning was associated with a nonsignificant 0.047 increase in AUC over regression (P = .134), external validation with a change of −0.039 (P = .132), and biomarker inclusion with −0.012 (P = .726). The advantage that biomarker and imaging models showed in individual studies, in other words, did not survive as a significant predictor once performance was pooled across the review. No pooled AUC was reported, because heterogeneity in model specification and follow-up duration made a summary estimate uninterpretable.

Where a Miscalibrated Model Does Damage

The stakes differ by where in the pathway a model sits. Screening-selection models shape who is invited for a scan, but the nodule-classification models feed directly into decisions about surveillance intervals, repeat imaging, PET, biopsy, or referral, and it was at that second stage that the authors located the sharpest risk of harm in either direction.

“Poor model performance at this stage can have far-reaching consequences, where underestimation may delay diagnosis of lung cancer, whereas overestimation may lead to unnecessary invasive procedures, complications, and patient distress,” the authors wrote.

The authors did not consider any of the reviewed models ready for clinical adoption. For eligibility, the tools already embedded in practice are PLCOm2012, the Bach and Liverpool Lung Project models, and the Lung Cancer Risk Assessment Tool. For nodules already detected, they are Lung-RADS, the Fleischner criteria, Mayo, and Brock/PanCan. Those remain the operative benchmarks, and validating or updating them should take precedence over building more, the authors argued.

The review carried limitations of its own. Heterogeneous reporting across the included studies, along with differing follow-up horizons, risk thresholds, and performance metrics, complicated direct comparison. Publication bias may favor models with higher apparent performance, and no formal risk-of-bias assessment was undertaken because the reported items were too incomplete and inconsistent across studies to rate consistently.

For any payer or health system being pitched a machine learning or biomarker-based risk tool, the review offers a compact test: a model earns its place only when it is well validated, well calibrated, and demonstrably better than what is already in use. The authors called for economic evaluation and implementation studies covering feasibility, cost-effectiveness, and measured performance under real clinical conditions before any of these tools enter routine screening workflows.

References

1. Rezaeianzadeh R, Leung C, Kim SJ, et al. Risk prediction for lung cancer screening: a systematic review and meta-regression. Eur Respir Rev. 2026;35(181):250295. doi:10.1183/16000617.0295-2025

2. Lung cancer: screening. US Preventive Services Task Force. March 9, 2021. Accessed August 27, 2026. https://www.uspreventiveservicestaskforce.org/uspstf/recommendation/lung-cancer-screening