Machine learning showed moderate external prediction of hypothyroidism after head and neck radiotherapy

Random forest reached external AUC 0.769 for 12-month radiation-induced hypothyroidism, but calibration and clinical utility weakened outside the development center.

KEY POINTS

  • This dual-center retrospective study included 256 patients in the Harbin development cohort and an independent 296-patient Kunming validation cohort. All patients received IMRT for head and neck cancers and had complete thyroid-function assessment through 12 months.
  • The primary endpoint was persistent radiation-induced hypothyroidism at 12 months, defined by TSH >4.340 mIU/L, including both clinical and subclinical hypothyroidism. Incidence was 21.1% (54/256) in the development cohort and 31.8% (94/296) externally.
  • Patients who developed hypothyroidism had smaller pretreatment thyroid volumes (11.72 vs 13.21 cm³, p < 0.001) and higher mean thyroid doses (52.40 vs 48.43 Gy, p = 0.009). Tumor site and clinical stage were also associated with the endpoint.
  • Six models were tested using 18 clinical and dosimetric features. Internally, LightGBM had the highest numerical performance with AUC 0.872, balanced accuracy 0.703, and PR-AUC 0.712. After correction for multiple comparisons, however, it was not significantly better than Random Forest, XGBoost, or SVM.
  • External performance dropped. Random Forest had the highest discrimination at AUC 0.769, with accuracy 72.3%, balanced accuracy 66.1%, and recall only 49.2%, meaning roughly half of hypothyroidism cases were missed at the selected threshold. SVM achieved higher recall at 71.2% but lower AUC at 0.702.
  • SHAP analysis identified pretreatment thyroid volume, T stage, tumor site, mean thyroid dose, and V50 among the strongest contributors. However, V30-V60 were highly intercorrelated (r = 0.75-0.98), and the authors explicitly caution against interpreting any individual Vx value as a fixed clinical dose threshold.
  • External calibration deteriorated, with models tending to underestimate risk at higher predicted probabilities. Decision-curve analysis also showed only limited external net benefit, and the authors conclude that the models are not yet suitable for clinical decision support. Follow-up was limited to 12 months, only 54 development events occurred, and patients dying or progressing before 12 months were excluded.

CLINICAL TAKEAWAY

The study reinforces that thyroid volume and overall thyroid dose matter, but it does not provide a new dose constraint. External validation exposed substantial loss of calibration and sensitivity, so the current models are best viewed as exploratory risk-stratification tools rather than something that should alter planning or follow-up today.

SOURCE

Cancers