MRI improved LLM-based recurrence prediction after definitive cervical chemoradiotherapy

Multimodal LLM prediction reached an AUC of 0.79, but AI-generated MRI reports showed clinically important discordances requiring expert review.

KEY POINTS

  • This retrospective single-centre exploratory study included 82 patients with cervical cancer treated with definitive radiotherapy between 2011 and 2016; 79 received concurrent platinum chemotherapy. During a median 5.0-year follow-up, 40 patients (48.8%) developed local or distant recurrence.
  • Treatment was relatively homogeneous: 92.7% received 50.4 Gy in 28 fractions of EBRT, 76.8% received 30 Gy in 5 brachytherapy fractions, and concurrent chemotherapy was delivered to 96.3%. Because completed chemotherapy cycles and overall treatment time were provided to the model, this was a retrospective treatment-complete risk assessment rather than true pretreatment prediction.
  • Gemini 3.1 Pro was tested without task-specific fine-tuning. Clinical and treatment data alone achieved an AUC of 0.704, which increased to 0.784 when the human MRI report was added and 0.762 when an AI-generated MRI report was substituted.
  • The best multimodal configuration — clinical information plus the AI-generated MRI report and a human lymph-node description — achieved an AUC of 0.790 versus 0.704 for the clinical baseline, an absolute gain of 0.085. This was significant on the prespecified Holm-corrected bootstrap analysis (p=0.032) but not confirmed by the Holm-corrected DeLong sensitivity analysis.
  • No significant difference was detected between predictions using human versus AI MRI reports (AUC 0.784 vs 0.762), but the study was neither designed nor powered to establish equivalence or non-inferiority.
  • AI-generated MRI reports remained the major weakness. Their mean rubric score against the clinical reference was 68.8%; T1/T2 and DWI descriptions agreed relatively well, but parametrial invasion had a 62% zero-score discordance rate. Lymph nodes were also frequently missed because the cropped tumor-centred images often excluded nodal stations.
  • Absolute risk estimates were poorly calibrated, with all strategies tending to underestimate recurrence risk (E/O 0.78–0.83). The cohort was small and single-centre, no full independent radiologist re-read of the underlying MRI was performed, and external or prospective validation is absent.

CLINICAL TAKEAWAY

The interesting signal here is not that an LLM can replace radiologists, but that a general-purpose model extracted additional prognostic information when MRI-derived information was added to the clinical record. With only 82 patients, inconsistent statistical confirmation and clinically meaningful MRI-report discordance, this remains a proof of concept rather than a deployable cervical-cancer risk model.

SOURCE

Cancers

Browse more research Suggest a correction