Authors
Yong Chen, Xiaoping Luo, Hao Zhang, Jia Zhou, Yonglin Yu, Guilan Yang, Juan Chen
Published in
Frontiers in endocrinology. Volume 17. Pages 1901646. Epub Aug 03, 2026.
Abstract
Atypical pulmonary carcinoid (AC) is a rare intermediate-grade neuroendocrine tumor with substantial clinical heterogeneity and an unpredictable prognosis. Accurate estimation of fixed-horizon mortality risk remains challenging because of its rarity and limited AC-specific prediction tools. This study aimed to develop and evaluate an interpretable model for estimating 3-year all-cause mortality in patients with AC.
Clinical data for patients with AC diagnosed between 2000 and 2021 were retrospectively obtained from the Surveillance, Epidemiology, and End Results (SEER) database. Patients diagnosed during 2000-2018 (n=1, 301) were randomly divided into a training set (n=910) and an internal test set (n=391). Patients diagnosed during 2019-2021 (n=446) constituted a temporal validation cohort within the same registry, and 45 patients treated at the General Hospital of Ningxia Medical University during 2015-2024 were used for preliminary independent single-center evaluation. Demographic, clinicopathological, and treatment variables were considered. LASSO regression and the Boruta algorithm selected 11 predictors: Grade, N_stage, M_stage, Bone_metastasis, Brain_metastasis, Liver_metastasis, Marital_status, Radiation, Chemotherapy, Age, and Tumor_Size. Seven machine learning models were developed to evaluate predictive performance, including Logistic Regression (LR), Extreme Gradient Boosting (XGBoost), Light Gradient Boosting Machine (LightGBM), Gradient Boosting Decision Tree (GBDT), Support Vector Machine (SVM), K-Nearest Neighbors (KNN), and Adaptive Boosting (AdaBoost). The SHapley Additive exPlanations (SHAP) approach was used to interpret feature importance.
LR showed the most consistent held-out performance, with an area under the receiver operating characteristic curve of 0.802 (95% CI: 0.751-0.852) in the internal test set, 0.839 (95% CI: 0.796-0.882) in the temporal validation cohort, and 0.904 (95% CI: 0.808-1.000) in the single-center cohort. Boosting models showed larger declines from apparent training to held-out performance. Decision curve analysis (DCA) suggested potential net benefit across selected threshold probabilities. SHAP identified Age as the largest contributor to model predictions.
The LR-based model showed consistent discrimination and interpretability in internal and same-registry temporal evaluation, with promising but imprecise results in the small single-center cohort. Geographic transportability remains unestablished. Independent multi-institutional validation with fully separated preprocessing is required before clinical use.
PMID:
42609244
Bibliographic data and abstract were imported from PubMed on 18 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 9
- Comments 0