Radiomics and Machine Learning in Medical Imaging / Ai in Cancer Detection / Thyroid Cancer Diagnosis and Treatment · Journal article
Annals of Medicine · September 10, 2026
Encouraging direction, but not yet definitive.
This is a machine learning classification model that segregates differentiated thyroid cancer patients into four post-I-131 response categories using pre-treatment biochemical markers and imaging features. The Random Forest model achieved AUC 0.894 in test data and 0.885 in independent temporal validation, identifying pre-stimulated thyroglobulin and sTg/TSH ratio as dominant predictors of excellent, indeterminate, and biochemical incomplete responses, while diagnostic whole-body scintigraphy was most predictive of structural incomplete response. The finding is promising but requires clinical outcome validation and comparison to existing risk stratification before adoption into practice.
Prospective machine learning model development with independent temporal validation cohort. Differentiated thyroid carcinoma patients with post-operative I-131 therapy. Specific eligibility criteria and clinical setting not reported.. Intervention: Machine learning model-based four-class response prediction (ER, IDR, BIR, SIR) using pre-treatment biochemical and imaging data. n = 948.
Random Forest achieved micro-average AUC 0.894 (95% CI: 0.842–0.901) in test cohort and 0.885 (95% CI: 0.832–0.912) in independent temporal validation cohort Pre-sTg dominated prediction of excellent response (ER), indeterminate response (IDR) and biochemical incomplete response (BIR), with peak weight of 31.9% in BIR Diagnostic whole-body scintigraphy (Dx-WBS) was definitive for structural incomplete response (SIR) prediction with 45.1% weight
Hard clinical outcomes (recurrence, disease-free survival, mortality) not reported; predictions are based on biochemical and imaging classification
Clinicians may consider this framework to stratify I-131 response risk before therapy; however, the model predicts response classification (a surrogate) rather than hard clinical outcomes. Validation against recurrence, mortality, or treatment modification effectiveness is needed before incorporating into standard pre-therapeutic decision-making.
Single-center machine learning model with reasonable test-set performance (AUC 0.894) and independent temporal validation, but surrogate endpoints (biochemical markers and imaging classification) rather than hard clinical outcomes, and no comparison to standard risk stratification.
As stated by the source record.
Quoted from the source exactly as published.
Clinicians may consider this framework to stratify I-131 response risk before therapy; however, the model predicts response classification (a surrogate) rather than hard clinical outcomes. Validation against recurrence, mortality, or treatment modification effectiveness is needed before incorporating into standard pre-therapeutic decision-making.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Background Differentiated thyroid carcinoma (DTC) responses to initial post-operative I-131 therapy are heterogeneous. Standard binary models combine incomplete responses into a ‘non-excellent’ category, obscuring transitional states (indeterminate response [IDR] and biochemical incomplete response [BIR]) and outcome-specific drivers. We developed and validated an interpretable four-class machine learning framework to separate these dynamic trajectories and support pre-therapeutic decisions.Materials and methods Data from 948 DTC patients were partitioned into training (n = 663) and test (n = 285) cohorts, with an independent temporal validation cohort (n=150). Four algorithms (Random Forest [RF], eXtreme Gradient Boosting [XGBoost], Light Gradient Boosting Machine [LightGBM] and Multi-Layer Perceptron [MLP]) were evaluated to predict excellent response (ER), IDR, BIR and structural incomplete response (SIR). Performance was assessed via area under the receiver operating characteristic curve (AUC), calibration curves and decision curve analysis (DCA). Shapley Additive exPlanations (SHAP) and multivariable logistic regression quantified category-specific drivers.Results Random forest achieved a micro-average AUC of 0.894 (95% CI: 0.842–0.901) in the test cohort and 0.885 (95% CI: 0.832–0.912) in validation. SHAP analysis revealed that pre-treatment stimulated thyroglobulin (pre-sTg) and stimulated thyroglobulin to thyroid-stimulating hormone ratio (LOG(sTg/TSH)) dominated ER, IDR and BIR predictions (pre-sTg peak weight: 31.9% in BIR), whereas diagnostic whole-body scintigraphy (Dx-WBS) was definitive for SIR (45.1% weight). Distant metastasis on Dx-WBS dramatically elevated SIR risk (OR: 171.89); pre-sTg ≥ 10 ng/mL significantly increased both SIR (OR: 7.18) and BIR (OR: 5.48) risks.Conclusions This four-class framework separates biochemical and structural drivers of post-I-131 responses, providing a practical risk-stratification tool for personalized management before therapy.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.