GPT-4.1 retrieved TrueBeam troubleshooting history but retained safety risks

A retrieval-augmented GPT-4.1 chatbot rapidly surfaced previous TrueBeam faults, but procedural, role-definition and safety errors remained.

KEY POINTS

  • The system indexed 1,394 troubleshooting logs collected from five TrueBeam linear accelerators across multiple hospitals over an eight-year period.
  • Each machine fault was stored as an independent structured text file within a retrieval-augmented GPT-4.1 environment to prevent unrelated events from being merged during retrieval.
  • The final dataset contained 5.44 MB of text. Indexing required 16.11 ± 0.52 minutes, supporting the feasibility of periodic re-indexing as new fault logs are added.
  • Mean response time across standardized questions was 7.50 ± 2.50 seconds, with a maximum observed response time of 14.44 seconds.
  • Four medical physicists evaluated responses covering historical recall, active troubleshooting, data aggregation, safety and temporal filtering.
  • The chatbot performed well when retrieving previous events, summarizing institutional experience and recognizing that information was unavailable in the indexed logs.
  • Important weaknesses included incorrect procedural order, inconsistent filtering by date, verbosity and unclear assignment of actions to physicists, service engineers or facilities personnel. Some responses proposed potentially unsafe steps without explicit escalation or patient-removal checks.

CLINICAL TAKEAWAY

Institution-specific retrieval tools could reduce the time required to locate prior fault histories and preserve operational knowledge. They should not independently direct machine recovery: explicit safety guardrails, role boundaries, source verification and formal failure-mode analysis are required before clinical deployment.

SOURCE

Technical Innovations & Patient Support in Radiation Oncology