KEY POINTS
- The systematic review included 40 studies published between 2015 and 2026, comprising 11 traditional metal-artifact reduction studies and 29 deep-learning studies focused on head-and-neck CT and radiotherapy.
- Deep-learning approaches ranged from CNNs and GANs to dual-domain networks, diffusion models and task-oriented architectures. However, only a minority had been tested against clinically meaningful RT endpoints rather than image-quality metrics alone.
- Across studies with clinical dosimetric validation, reported dose deviations were generally <3%**, while several studies achieved gamma pass rates **>95% at 3%/3 mm.
- One task-oriented masked-MSE model achieved a mean organ-at-risk dose error of 0.8%, gamma pass rates >95% and a PTV D98 difference of ≤0.3 Gy, illustrating the potential benefit of training directly toward dosimetric rather than purely visual accuracy.
- GAN-based methods have the strongest current clinical dosimetric evidence, with some studies reporting dose errors below 1%. However, image-domain generative models remain vulnerable to hallucinated anatomy and domain shift, particularly when trained on synthetic artifacts.
- Dual-domain and diffusion-based approaches achieved excellent image-quality metrics but generally lacked direct radiotherapy dosimetric validation, so their apparent technical superiority cannot yet be assumed to translate into safer treatment planning.
- No included study qualified as prospective Level 1 clinical validation. Most models were trained or tested on synthetic or single-centre datasets, raw sinogram access remains difficult, and prospective multicentre validation is still absent.
CLINICAL TAKEAWAY
Deep-learning metal-artifact reduction is reaching dosimetric performance that could be useful in head-and-neck RT, particularly when the target lies close to dental hardware. But better-looking CT images are not enough: local validation should focus on dose calculation, target coverage and failure modes before these tools are trusted clinically.