Deep learning reduced head and neck auto-segmentation editing, but user variability persisted

Deep learning auto-segmentation reduced editing, but both minimal and excessive manual edits were associated with lower clinical acceptability.

KEY POINTS

  • This single-institution retrospective study analyzed 1051 head and neck patients treated between 2017 and 2024, comparing auto-segmentations with final clinical contours for the parotid glands, submandibular glands, oral cavity, and glottis.
  • The study covered several workflow eras: fully manual delineation, 11-atlas auto-segmentation, 22-atlas auto-segmentation, and deep learning-based auto-segmentation introduced in January 2023.
  • After initial atlas-based auto-segmentation implementation, median mean surface distance decreased for all evaluated structures except the glottis; for the left parotid, it decreased from 2.0 mm to 1.3 mm.
  • Deep learning markedly reduced editing across structures; for the left parotid, mean normalized added path length decreased from 0.70 before deep learning to 0.25 afterward.
  • Editing behavior varied substantially between radiation therapy technologists, and both minimal and extensive editing were associated with lower peer-rated clinical acceptability.

CLINICAL TAKEAWAY

Auto-segmentation quality is only part of the clinical workflow problem: user behavior after deployment can still determine whether efficiency and contour consistency improve. This study supports routine per-user monitoring of editing magnitude and editing locations, especially after software upgrades. The evidence is limited by its single-institution design, lack of consensus ground truth, small acceptability subset, and absence of dosimetric impact analysis.

SOURCE

Radiotherapy and Oncology