Deep learning generated deliverable lymphoma intensity-modulated plans but remained dosimetrically inferior

Deep learning produced machine-deliverable lymphoma plans across heterogeneous anatomy, but target coverage and hot spots remained inferior to reference planning.

KEY POINTS

  • The study included 571 consecutive lymphoma patients and 619 clinical plans, spanning curative and palliative treatment, targets from 5.4 to 2,282 cm³, and prescriptions from 4 to 60 Gy. Clinical volumetric modulated arc plans were automatically converted into standardized 15-field intensity-modulated reference plans, yielding 9,285 beam-level training pairs.
  • The main single-prescription dataset contained 302 plans, split at patient level into 195 training, 46 validation and 61 held-out test plans. Another 317 multi-prescription plans were used for pretraining to expose the models to broader dose patterns.
  • The reference intensity-modulated replanning itself produced slightly colder plans than the original clinical arc plans: mean target D98% was 92.71% versus 94.37%, D2% 102.24% versus 104.41%, and conformity index 0.781 versus 0.877.
  • Models incorporating a dose-volume-histogram clinical loss achieved the best intermediate dose predictions. Repeated training of one such configuration produced target D98% of approximately 98.0±0.5%, demonstrating that optimizing clinically meaningful endpoints can outperform simple voxel-by-voxel reference replication at the prediction stage.
  • The dose-to-fluence network achieved mean absolute error 0.015±0.009 against reference fluence maps. After import into Eclipse and multileaf-collimator sequencing, predicted and sequenced fluences remained closely matched, with 95.8±2.6% gamma passing at 3%/3 mm and no beam failing sequencing.
  • Despite successful machine sequencing, the final end-to-end plans lost quality during the transition from predicted dose to deliverable fluence. Even the best configuration was significantly inferior to the intensity-modulated reference on all reported target metrics (all p<0.001), with lower coverage and higher hot spots.
  • The model used computed tomography and target information without explicit organ-at-risk contours, was trained within one institution and one treatment-planning ecosystem, and was not subjected to clinical plan acceptance. The authors therefore do not support autonomous clinical use.

CLINICAL TAKEAWAY

The difficult part is no longer merely generating a fluence map that a machine can deliver. This study shows that an artificial intelligence model can close the technical loop through a commercial treatment-planning system, but preserving clinically competitive dose quality through that loop remains the unresolved problem.

SOURCE

Journal of Applied Clinical Medical Physics