KEY POINTS
- The longitudinal analysis compared three clinical eras using 150 rectal cancer patients, with 50 each before AI implementation, after first-generation deployment and after second-generation deployment. An additional 21 cases were contoured by six radiation oncologists using manual, Auto1-assisted and Auto2-assisted workflows.
- The first system, introduced in 2018, used DeepLab V3; the 2024 update replaced it with nnU-Net V2. In routine use, AI contours were transferred into the treatment-planning system and then reviewed and edited by radiation oncologists rather than accepted automatically.
- After the system update, agreement between unedited AI contours and clinically approved contours improved substantially: mean CTV Dice increased from 0.88 to 0.93 (P<.001) and mean OAR Dice from 0.88 to 0.95 (P<.001).
- Mean contouring failure rate fell from 3.30% with the first-generation model to 0.64% with the second generation, an approximately 80.6% relative reduction. Individual second-generation failure rates ranged from 0% for colon to 1.67% for femoral structures.
- Workflow gains were large. Mean total contouring time was 41.15 minutes manually, 21.74 minutes with Auto1 and 16.97 minutes with Auto2, meaning the updated system reduced total time by 58.8% versus manual contouring and 21.9% versus Auto1.
- Auto2 also improved reproducibility. CTV interobserver Dice was 0.95 with Auto2, 0.93 with Auto1 and 0.92 manually, while CTV accuracy against the expert-derived reference was 0.94, 0.93 and 0.91, respectively. Similar gains were seen for several OARs.
- Blinded senior review found 99.2% of oncologist-finalized contours clinically acceptable, while unedited Auto2 contours scored 4.02 versus 3.26 with Auto1 on a five-point scale (P<.001). The study nevertheless represents one institution, one disease site and locally developed models, so the magnitude of benefit may not transfer directly elsewhere.
CLINICAL TAKEAWAY
The important finding is not simply that AI contours were faster: updating the deployed model materially improved efficiency, consistency and failure rates after years of real-world use. It supports continuous QA and model lifecycle management rather than treating auto-contouring software as a one-time installation, while retaining physician review.