KEY POINTS
- This structured mini-review searched PubMed, Ovid MEDLINE and Embase through 17 March 2026 and identified 10 contemporary studies using multimodal deep learning to predict radiotherapy-related toxicity in head and neck cancer. It was not a formal systematic review, had no registered protocol, and screening/data extraction were performed by a single reviewer.
- The reviewed models combined clinical variables with spatial CT, dose distributions, organ segmentations or longitudinal imaging, using early, intermediate or late fusion architectures. Fusion strategy was often poorly reported, despite evidence that different fusion approaches can materially change discrimination and calibration.
- Performance gains over conventional models were inconsistent. For 12-month xerostomia, one early study reported AUC 0.84 versus 0.74 with clinical logistic regression; however, a later externally validated model reached only AUC 0.66 after transfer learning versus 0.63 when trained de novo, and its reference NTCP model initially performed better externally.
- Dysphagia prediction showed one of the stronger external signals. A multimodal model integrating clinical data, CT, dose and segmentations achieved external AUC 0.74 versus 0.63 for a conventional NTCP model, while intermediate fusion provided the best calibration.
- More complexity did not consistently improve prediction. In a 1,012-patient study, adding imaging and dosimetric information did not improve feeding-tube prediction over clinical features alone; another multi-toxicity model improved external dysphagia prediction to AUC 0.71, yet conventional logistic regression remained better for aspiration at AUC 0.72 versus 0.68 with deep learning.
- External validation was uncommon and frequently exposed performance loss. The review also found inconsistent reporting of missing-data handling, hyperparameter tuning, model selection and calibration; Figure 1 shows that discrimination metrics were nearly universal, whereas calibration, decision-curve analysis and external validation were much less consistently reported.
- None of the reviewed contemporary multimodal models integrated the full combination of clinical, imaging, dosimetric and biological/genomic data, and none demonstrated sufficient prospective clinical utility for routine treatment planning. The authors identify larger multicenter datasets, harmonized toxicity endpoints, external validation, calibration, reproducibility and prospective decision-impact testing as prerequisites for deployment.
CLINICAL TAKEAWAY
Multimodal deep learning can sometimes outperform traditional toxicity models in head and neck radiotherapy, but the benefit is endpoint- and dataset-dependent rather than consistent. For now, these models remain research tools: strong discrimination alone is insufficient without reliable calibration, external validation and evidence that predictions actually improve treatment decisions or patient outcomes.