KEY POINTS
- The retrospective single-center cohort included 88 patients with unresectable stage III-IV head and neck squamous cell carcinoma treated with definitive platinum-based chemoradiotherapy to a median EQD2 of 70.7 Gy. Complete response occurred in 67 patients, while 21 had a non-complete response at 3-6 months.
- Clinical predictors were already strong. Non-responders had substantially larger primary tumors, with median GTV 66.8 versus 20.7 cm³ (p=0.006), and primary tumor site was strongly associated with response (p<0.001). Oral cavity tumors had the highest non-response rate at 64.3%.
- The investigators evaluated 366 machine-learning configurations across clinical variables, CT radiomics, multiparametric MRI radiomics and multiple early- and late-fusion strategies using nested leave-one-out cross-validation.
- The single best MRI pipeline produced an impressive AUC of 0.91, with 90% sensitivity and 88% specificity. However, it was not significantly better than the best clinical model, which achieved AUC 0.75 (DeLong p=0.058).
- Looking across all configurations rather than selecting only the winner changed the interpretation substantially. Mean AUC was 0.683 for clinical data, 0.614 for CT and only 0.548 for MRI. MRI therefore averaged 0.135 AUC lower than the clinical baseline across configurations.
- Imaging added consistent value only when combined with clinical information at the decision level. Late fusion of clinical and CT models increased mean AUC by 0.059, with a patient-level bootstrap 95% CI of +0.007 to +0.118. CT plus MRI without clinical data provided no average improvement.
- The apparent MRI signature also correlated strongly with known clinical factors such as primary site and T stage, suggesting that radiomics partly re-encoded information already available clinically. With only 21 non-response events, no external cohort and a single MRI platform, treatment intensification based on these models would be premature.
CLINICAL TAKEAWAY
The most important finding is methodological rather than predictive: reporting only the best radiomics pipeline can make performance look much stronger than it really is. In this dataset, conventional clinical information remained remarkably competitive, and radiomics should not yet be used to select patients for dose escalation or treatment intensification.