KEY POINTS
- A hierarchically densely connected U-Net was developed using clinical head and neck intensity-modulated radiotherapy plans from one Danish institution: 388 patients were used for training and validation, and 42 formed an internal test cohort.
- External validation used 560 clinically delivered plans from six Danish centres in the national DAHANCA repository. The cohort included oral cavity, oropharyngeal, hypopharyngeal and laryngeal cancers planned in Eclipse or Pinnacle.
- Treatments used simultaneous integrated boost prescriptions of 66 Gy in 33 fractions or 68 Gy in 34 fractions to high-dose targets, with 60 Gy to intermediate-risk and 50 Gy to elective volumes.
- Across the evaluation cohorts, the median predicted-minus-clinical mean-dose difference was -0.1 Gy for planning target volumes, with an interquartile range of -0.5 to 0.3 Gy. For organs at risk, the median difference was 1.1 Gy, with an interquartile range of -0.6 to 3.7 Gy.
- Target coverage remained close to the clinical plans: absolute median differences in clinical target volume D99% and planning target volume D98% were within 0.4 Gy internally and 1.1 Gy in the multicentre cohort.
- The deep-learning model outperformed a simple comparator assigning median training-cohort doses. The comparator produced target and organ-at-risk differences of 0.7 Gy and -4.6 Gy, respectively, with organ-at-risk interquartile ranges approximately four to five times wider.
- In the external cohort, predicted-minus-clinical differences were 4.4 percentage points for the normalized toxicity index, 1.2 points for grade 2 or higher xerostomia probability and 1.1 points for grade 2 or higher dysphagia probability. These differences could provide patient-specific thresholds for plan-review alerts.
- Ninety-four external patients had missing organ-at-risk contours that required artificial-intelligence-assisted recontouring and quality assurance. Centre-specific contouring, optimization priorities and planning systems remained potential sources of prediction error.
CLINICAL TAKEAWAY
Single-centre deep-learning dose prediction remained reasonably accurate across a heterogeneous national cohort and could provide a patient-specific benchmark during plan review. It should currently be treated as a screening tool rather than an autonomous judge of plan quality, because predictions were compared with historical clinical plans rather than prospectively optimized alternatives.