KEY POINTS
- The study evaluated 28 months of routine clinical deep learning-based IMPT planning for oropharyngeal cancer at UMC Groningen. Of 99 patients considered from March 2023 to June 2025, 78 ultimately received a deep learning-generated proton plan.
- Prescription was 70 Gy(RBE) to the therapeutic CTV and 54.25 Gy(RBE) to bilateral elective nodal volumes. A 3D U-Net predicted voxel-wise dose, followed by robust dose-mimicking optimization using 3-mm setup and 3% density uncertainty and Monte Carlo dose calculation.
- The initial automated plan was not simply accepted unchecked. Treatment planners manually refined it, generally requiring around 2 hours, then compared OAR doses with an independent prediction tool; final plans underwent multidisciplinary slice-by-slice review by a radiation oncologist, planner and medical physicist.
- Manual refinement increased achievement of therapeutic CTV goals from 91.0% to 93.6% and elective CTV goals from 78.2% to 88.5%. Predicted grade ≥2 dysphagia fell from 15.9% to 15.5% and xerostomia from 37.9% to 37.6%; these changes were statistically significant but clinically small.
- All 78 final deep learning plans were considered clinically acceptable and used for treatment. Sixty-eight met every predefined clinical goal; in the remaining 10, CTV V94% was narrowly below the 98% criterion at 97.5–97.9%.
- Every fifth case generated a blinded manual benchmark, providing 16 paired comparisons. Manual plans met all goals in 16/16 versus 14/16 deep-learning plans, but deep-learning plans reduced spinal cord D0 from 33.2 to 30.3 Gy(RBE) (p=0.004) and mean parotid dose from 16.9 to 15.8 Gy(RBE) (p=0.002), with no other significant differences.
- The independent OAR prediction tool correlated very strongly with achieved doses (r=0.96–1.00). However, the study came from one experienced center, did not quantify actual workflow time prospectively and retained substantial human oversight, so it does not demonstrate safe autonomous planning.
CLINICAL TAKEAWAY
This is an important example of what real-world AI implementation looks like: automation combined with independent dose guidance, manual refinement and multidisciplinary review rather than a fully autonomous planning system. The technology performed well for oropharyngeal IMPT, but external validation and explicit workflow-efficiency data remain necessary.