KEY POINTS
- The controlled experiment recruited 36 medically trained radiology professionals at ECR 2025 to inspect and manually delineate pediatric diffuse midline gliomas on T2-FLAIR MRI while eye and mouse movements were recorded. After exclusions, the final dataset contained 28 participants, 110 trials, and 20,238 contour points.
- The image dataset comprised 12 two-dimensional axial slices from 10 diffuse midline glioma patients. Interobserver contour distance was used as the proxy for local uncertainty, while image-derived features included U-Net segmentation probabilities, Grad-CAM saliency and local image entropy.
- Human segmentation performance correlated with U-Net performance (|r|=0.40, p<0.001), while local segmentation error correlated more strongly with local uncertainty (|r|=0.60). Even when the U-Net failed completely, the model estimated that human observers would still delineate approximately half of the tumor.
- A mixed-effects model combining behavioral and image features explained 30% of uncertainty through fixed effects. Image-derived features alone explained substantially more than behavioral features alone (marginal R² 0.24 versus 0.02), although participant- and scan-specific effects remained large.
- Greater uncertainty was associated with faster saccades and greater fixation density, while longer fixations during verbal assessment were associated with lower uncertainty. High-performing annotators generally used fewer fixations and saccades, shorter gaze paths and more spatially focused attention during contouring.
- A random-forest model evaluated on unseen participants and scans explained 39.2% of uncertainty variance, with MAE approximately 0.2. U-Net-derived features and image entropy dominated feature importance, while faster mouse movement and higher saccade velocity were the strongest behavioral indicators.
- There is a reporting inconsistency worth noting: the abstract states that downsampling-layer features were the strongest predictors, whereas the Results section identifies saliency from the third and fourth upsampling layers plus image entropy among the leading random-forest predictors. The experiment also used single low-resolution 2D slices without scrolling, zooming or multimodal imaging.
CLINICAL TAKEAWAY
Passive eye and cursor tracking could eventually add an uncertainty layer to manually generated contours, highlighting regions that deserve review or different margins. The current study demonstrates feasibility rather than a usable clinical QA system: full 3D contouring, specialist radiation oncologists, routine software and dose-impact validation are still required.