Uncertainty estimates flagged neural-network dose modelling failures in MR-guided radiotherapy

Calibrated uncertainty thresholds kept gamma-defined dose-modelling failure below 10% while accepting 45–77% of irradiation segments.

KEY POINTS

  • The study extended a previously validated 3D U-Net dose-modelling framework using 6,713 irradiation segments from 130 patients treated on a 1.5-T MR-Linac. The dataset included prostate, liver, head and neck, partial-breast and nodal treatments; full EGSnrc Monte Carlo simulations served as ground truth.
  • Three uncertainty strategies were compared: Monte Carlo dropout, mean variance estimation and a six-model deep ensemble. The models retained strong baseline dose accuracy, with mean 3%/3-mm gamma pass rates of 97.4%, 97.8% and 97.6%, respectively.
  • Calibration markedly improved the reliability of predicted uncertainty. Expected normalized calibration error fell to 0.14 with Monte Carlo dropout, 0.05 with mean variance estimation and 0.10 with deep ensembles, making mean variance estimation the best-calibrated approach.
  • Mean predicted uncertainty tracked actual dose error strongly: Spearman correlation with mean absolute error was 0.64, 0.76 and 0.67, respectively. Mean variance estimation therefore showed the strongest relationship between uncertainty and true prediction error.
  • The authors then defined uncertainty thresholds targeting a 10% failure rate, with failure defined as 3%/3-mm gamma passing rate <95%. In independent evaluation, failure among accepted segments was 7.2% with Monte Carlo dropout, 7.5% with mean variance estimation and 9.1% with deep ensembles.
  • The practical trade-off differed substantially. Monte Carlo dropout accepted only 45% of segments but had 0.77 sensitivity for identifying failures; mean variance estimation accepted 60%, while the deep ensemble accepted 77% and had the highest specificity at 0.82, though sensitivity fell to 0.49.
  • Computational burden favored mean variance estimation, which required only one network and one forward pass. Deep ensembles and Monte Carlo dropout required approximately 6-fold and 10-fold greater inference cost, respectively; prostate plans showed the lowest uncertainty and head-and-neck plans the highest.

CLINICAL TAKEAWAY

AI dose calculation becomes much more clinically useful if the system can also identify when its own prediction is unreliable. These uncertainty estimates could route high-risk segments back to Monte Carlo calculation while retaining faster neural-network calculation for the rest, but the thresholds still require prospective validation in the intended clinical workflow.

SOURCE

Physics and Imaging in Radiation Oncology