Audit logs measured artificial intelligence contour correction time with 0.93 concordance

Treatment-planning audit logs closely reproduced manual contour-edit timing and exposed major structure-specific differences in correction workload.

KEY POINTS

  • The retrospective clinical dataset covered 1,967 patients and 24,970 artificial intelligence-generated structures reviewed in RayStation between March 2023 and June 2025. Of these, 3,083 structures underwent manual correction by 23 radiation oncologists.
  • Audit-log timestamps were converted into active editing time using a maximum 60-second interval between consecutive structure-editing actions. Against manually recorded timing for 54 structures from eight cases, the automated method achieved Lin's concordance correlation coefficient of 0.93.
  • Only 12% of all generated structures were edited. Brush tools were used in 80% of edited contours, interpolation in 42%, smart brush in 15%, three-dimensional deformation in 6%, and spline drawing in 4%.
  • Interpolation was associated with larger corrections and longer editing: median time was 81 seconds versus 58 seconds without interpolation, while added path length was 232 mm versus 18 mm. The correlation between geometric change and editing time was only fair (r=0.38 with interpolation; r=0.51 without).
  • Workload differed dramatically by structure. The mandible had a 52% correction rate and median edit time of 374 seconds, whereas left and right lung correction rates were only 1%; median cochlear edits took approximately 13–16 seconds.
  • The relationship between geometry and time was highly structure- and tool-dependent. For example, bladder correction time correlated moderately with added path length overall (r=0.64) but only fairly when interpolation was used (r=0.47).
  • The approach was developed in one commercial treatment-planning environment, and audit logs estimate active interaction rather than every component of clinical review. Cross-platform standardization will therefore be necessary before comparing artificial intelligence systems or institutions using this metric.

CLINICAL TAKEAWAY

Dice scores alone cannot tell a department how much work an auto-contouring model actually saves. Audit-log-derived edit time offers a low-burden way to monitor real clinical workload continuously and may be a more meaningful operational endpoint when commissioning or updating segmentation models.

SOURCE

Physics and Imaging in Radiation Oncology