
Simplest Model Won AI Comparison for Autoimmune Hepatitis Recurrence
Key Takeaways
- Reported AIH recurrence rates after transplant vary from 17%–42%, largely driven by protocol biopsy practices and follow-up duration, with recurrence risking fibrosis, graft loss, and reduced survival.
- In 706 recipients across 33 centers, logistic regression achieved the best discrimination and calibration versus tree-based machine learning models, underscoring limited incremental value from algorithmic complexity.
Logistic regression beats machine learning to predict autoimmune hepatitis recurrence after liver transplantation.
Three machine learning models trained to predict recurrent autoimmune hepatitis after liver transplantation were all outperformed by ordinary logistic regression, which reached an area under the receiver operating characteristic curve (AUROC) of 0.747 against 0.730, 0.674, and 0.553 for random forest, XGBoost, and gradient boosting, respectively. The comparison ran across 706 recipients treated at 33 centers on 4 continents.1
A Recurrence Rate Nobody Agrees On
The immunosuppressive regimen that best prevents it remains unsettled. Long-term low-dose corticosteroids have been proposed as protective, while American Association for the Study of Liver Diseases guidance holds that glucocorticoids can be stopped after transplant with monitoring for recurrence. Mycophenolate mofetil in the first posttransplant year has been linked to roughly 3-fold higher recurrence risk in earlier work, and azathioprine to lower risk, although azathioprine's advantage is confounded by the 1980s and 1990s era in which it was mostly used.
706 Recipients, 62 Variables
The cohort drew on 855 patients transplanted for AIH between 1987 and 2020, excluding 119 with overlap syndrome and 30 who recurred within the first year, leaving 706. Recurrence was defined histologically, requiring interface with hepatitis and a biopsy negative for T-cell–mediated rejection. Females made up 76.6%, the mean age at transplant was 41.9 years, and 117 patients (16.6%) recurred beyond 1 year.
Sixty-two features entered model training, filtered by a missingness cutoff below 20%, covering demographics; immunologic markers; serial liver chemistries at transplant and 3, 6, and 12 months after; donor variables; explant histology; rejection episodes; and immunosuppressive drugs and trough levels. Delta variables captured laboratory trajectories rather than single points. The data were split 70% for training and 30% for internal validation, with performance assessed across 1000 bootstrap iterations and reporting following TRIPOD+AI, the standard for machine learning prediction models.
The Simplest Model Won
Logistic regression produced the best discrimination at an AUROC of 0.747 (95% CI, 0.679-0.815), with a sensitivity of 0.655 and a specificity of 0.720. Random forest reached 0.730, XGBoost 0.674, and gradient boosting 0.553, the last with a sensitivity of 0.241 and a strong bias toward predicting no recurrence. Calibration of the regression model held across the risk range, with a Brier score of 0.108, meaning predicted probabilities tracked observed recurrence rates closely.
Errors landed in a consequential pattern. Of 212 patients in the test set, 41 were classified as high risk without recurring, and 5 who did recur were missed. Escalating immunosuppression in the first group invites infection, metabolic complications, and malignancy; missing the second permits graft injury.
Predicted probabilities spread across the range rather than clustering near the midpoint, with 2.8% of patients below 0.2, 40.6% between 0.2 and 0.4, 43.9% between 0.4 and 0.6, and 12.7% between 0.6 and 0.8. No patient exceeded 0.8, a fair description of how much certainty the available variables can produce.
That gap is not unique to this model. A systematic review of 65 artificial intelligence studies in liver transplantation found algorithms that often outperformed established risk scores such as the Model for End-Stage Liver Disease, but also insufficient external validation and risk of bias across the literature, with prospective validation still a prerequisite for adoption.2
What the Model Says About Drugs
Shapley additive explanations ranked a 1-year tacrolimus trough at or above 6 ng/mL as the strongest contributor to predicted recurrence, followed by recipient-donor sex mismatch, rejection history, younger age at transplant and diagnosis, elevated creatinine and international normalized ratio before transplant, and higher imunoglobulin G.1 Among 245 patients with a trough above 6 ng/mL, 21.2% recurred.
The investigators were careful about what that does not mean. Because first-year recurrence was excluded, they read the 1-year trough as a marker of patients whose immunologic activity had already prompted more immunosuppression, not as evidence that higher exposure causes recurrence. Initial tacrolimus use pointed the other way and appeared protective, with the drug used in 59.0% of recurrence cases vs 81.8% of the rest, while cyclosporine appeared in 24.3% vs 8.3%. Adding long-term prednisone to tacrolimus and mycophenolate mofetil conferred no further protection.
Speaking to the limitations of their investigation, the authors highlighted that a retrospective design cannot establish causality, immunosuppressive choices reflect physician judgment as much as drug properties, and follow-up was too inconsistent to evaluate graft fibrosis or retransplantation. They positioned the output as a risk-stratification aid for monitoring intensity and shared decision-making, not a trigger for automatic regimen changes, and called for prospective validation first. That framing matters, because an AUROC of 0.747 from the least elaborate model in the comparison is a modest result wearing the label of artificial intelligence.
References
1. Bhat M, Sun Y, Manickavel P, et al. Artificial intelligence predicts recurrent autoimmune hepatitis after liver transplantation in a multicenter cohort study. Hepatol Commun. 2026;10(8):e1004. doi:10.1097/HC9.0000000000001004
2. Boutos P, Rogers J, Kouvela E, Voulgaris G, Thaker S, Tsoulfas G. f. J Surg Res. 2026;326:786-793. doi:10.1016/j.jss.2026.07.039
Related to this article








