KEY POINTS
- This single-institution retrospective study analyzed 1051 head and neck patients treated between 2017 and 2024, comparing auto-segmentations with final clinical contours for the parotid glands, submandibular glands, oral cavity, and glottis.
- The study covered several workflow eras: fully manual delineation, 11-atlas auto-segmentation, 22-atlas auto-segmentation, and deep learning-based auto-segmentation introduced in January 2023.
- After initial atlas-based auto-segmentation implementation, median mean surface distance decreased for all evaluated structures except the glottis; for the left parotid, it decreased from 2.0 mm to 1.3 mm.
- Deep learning markedly reduced editing across structures; for the left parotid, mean normalized added path length decreased from 0.70 before deep learning to 0.25 afterward.
- Editing behavior varied substantially between radiation therapy technologists, and both minimal and extensive editing were associated with lower peer-rated clinical acceptability.
CLINICAL TAKEAWAY
Auto-segmentation quality is only part of the clinical workflow problem: user behavior after deployment can still determine whether efficiency and contour consistency improve. This study supports routine per-user monitoring of editing magnitude and editing locations, especially after software upgrades. The evidence is limited by its single-institution design, lack of consensus ground truth, small acceptability subset, and absence of dosimetric impact analysis.