KEY POINTS
- Three commercial deep-learning contouring systems—Mirada DLCExpert, Limbus AI and MVision AI Contour+—were evaluated in 20 patients with laryngeal or hypopharyngeal cancer. Seven OARs and cervical nodal levels I–V were compared with contours from an experienced radiation oncologist.
- Pooled across seven OARs, mean Dice scores were 0.80 for Mirada, 0.85 for Limbus and 0.85 for MVision. The apparently poorer Mirada result was almost entirely driven by the esophagus and spinal canal, where the vendor used substantially different anatomical definitions.
- For the five OARs where definitions were comparable—brainstem, bilateral parotids and bilateral submandibular glands—Mirada actually had the highest mean Dice score: 0.92 versus 0.84 with Limbus and 0.85 with MVision. Pairwise advantages became much more limited after correction for multiple comparisons.
- The definition mismatch was dramatic for two structures. Mirada achieved mean Dice scores of only 0.40 for the esophagus and 0.57 for the spinal canal, with corresponding volumes only 35% and 44% of the expert reference. The authors interpret this primarily as a difference in what the vendor calls the structure rather than a conventional segmentation failure.
- For merged cervical nodal levels, mean Dice scores were 0.84, 0.82 and 0.79 for Mirada, Limbus and MVision. MVision had the narrowest variability and was the only model without systematic volume bias versus the expert reference (+3.8 cm³), whereas Mirada underestimated volume by 33.4 cm³ and Limbus overestimated by 30.5 cm³.
- Clinical review showed an important distinction between geometric scores and usability. Major corrections were required in 25.0% of Mirada assessments versus 6.4% for Limbus and 5.4% for MVision, while contours were accepted without modification in 27.9%, 50.0% and 48.6%, respectively.
- The workflow gain was substantial: fully manual contouring averaged 28 minutes for OARs plus 16 minutes for nodal levels—44 minutes total—versus approximately 5 minutes for AI-assisted review. However, the study included only 20 patients, part of the reference dataset was not independent of the tested models, and no dosimetric impact analysis was performed.
CLINICAL TAKEAWAY
Commercial head-and-neck auto-contouring can transform workflow time, but a high Dice score is not enough to commission a model safely. The most useful lesson is that vendor-specific anatomical definitions must be checked structure by structure, particularly for the esophagus, spinal canal and nodal volumes, before automated contours enter routine planning.
SOURCE
Technical Innovations & Patient Support in Radiation Oncology