TY  - JOUR
A1  - Riepe, Christoph
A1  - van de Water, Robin
A1  - Winter, Axel
A1  - Pfitzner, Bjarne
A1  - Faraj, Lara
A1  - Ahlborn, Robert
A1  - Schulze, Maximilian
A1  - Zuluaga, Daniela
A1  - Schineis, Christian
A1  - Beyer, Katharina
A1  - Pratschke, Johann
A1  - Arnrich, Bert
A1  - Sauer, Igor M.
A1  - Maurer, Max M.
T1  - 90-day mortality prediction in elective visceral surgery using machine learning: a retrospective multicenter development, validation, and comparison study
T2  - International Journal of Surgery
N2  - Background: 

Machine Learning (ML) is increasingly being adopted in biomedical research, however, its potential for outcome prediction in visceral surgery remains uncertain. This study compares the potential of ML methods for preoperative 90-day mortality (90DM) prediction of an aggregated multi-organ approach to conventional scoring systems and individual organ models.

Methods: 

This retrospective cohort study enrolled patients undergoing major elective visceral surgery between 2014 and 2022 across two tertiary centers. Multiple ML models for preoperative 90DM prediction were trained, externally validated and benchmarked against the American Society of Anesthesiologists (ASA) score and revised Charlson Comorbidity Index (rCCI). Areas under the receiver operating characteristic (AUROC) and precision recall curves (AUPRC) including standard deviations were calculated. Additionally, individual models for esophageal, gastric, intestinal, liver, and pancreatic surgery were developed and compared to an aggregated approach.

Results: 

7711 cases encompassing 78 features were included. Overall 90DM was 4% (n = 309). An XBoost classifier demonstrated the best performance and high robustness following external validation (AUROC: 0.86 [0.01]; AUPRC: 0.2 [0.04]). All models outperformed the ASA score (AUROC: 0.72; AUPRC: 0.08) and rCCI (AUROC: 0.81; AUPRC: 0.11). rCCI, patient age and C-reactive protein emerged as most decisive model weights. Models for gastric (AUROC: 0.88 [0.13]; AUPRC: 0.24 [0.26]) and intestinal surgery (AUROC: 0.87 [0.05]; AUPRC: 0.17 [0.09]) revealed the highest organ-specific performances, while pancreatic surgery yielded the lowest results (AUROC: 0.66 [0.08]; AUPRC: 0.22 [0.12]). A combined multi-organ approach (AUROC: 0.84 [0.04]; AUPRC: 0.21 [0.06]) demonstrated superiority over the weighted average across all organ-specific models (AUROC: 0.82 [0.07]; AUPRC: 0.2 [0.13]).

Conclusion: 

ML offers robust preoperative risk stratification for 90DM in elective visceral surgery. Leveraging training across multi-organ cohorts may improve accuracy and robustness compared to organ-specific models. Prospective studies are needed to confirm the potential of ML in surgical outcome prediction.
Y1  - 2025
UR  - https://opus.bibliothek.uni-augsburg.de/opus4/frontdoor/index/index/docId/123704
UR  - https://nbn-resolving.org/urn:nbn:de:bvb:384-opus4-1237044
SN  - 1743-9159
VL  - 111
IS  - 6
SP  - 3742
EP  - 3751
PB  - Ovid Technologies (Wolters Kluwer Health)
ER  -