KEY POINTS
- The framework was tested in three lung cancer patients with ten-phase 4DCT datasets and peak superior–inferior tumour motion of 0.7, 2.9, and 10.6 mm, respectively. All plans prescribed 60 Gy in 30 fractions.
- Separate deep Q-learning agents were trained for each patient and proton energy layer using only the mid-position CT. During inference, the agents used two-dimensional target projections, beam position, and prior spot delivery to select the next beam position and deliver spots in 0.5-MU increments, without retraining on individual respiratory phases.
- Comparison plans included a static GTV plan optimized on the mid-position CT and an ITV plan encompassing the GTV across all ten respiratory phases. Every plan used a single proton field at a 270° gantry angle.
- Compared with the static GTV plans, reinforcement learning increased mean GTV D95 by 4.54 Gy, 2.20 Gy, and 0.20 Gy in the three patients. The benefit was smallest in the patient with the largest respiratory displacement.
- Compared with ITV planning, reinforcement learning reduced mean lung-minus-GTV dose by 0.42, 4.12, and 3.13 Gy, respectively, and reduced mean heart dose by 1.41 Gy in the third patient. These gains came with lower GTV D95 by 2.73–6.99 Gy and less homogeneous target dose than ITV planning.
- Some constraints were exceeded: lung-minus-GTV V20 reached 30.46% in the third patient, and three respiratory phases in the first patient exceeded the spinal-cord D0.001cm³ limit. Each decision step required approximately 2.6 milliseconds.
- Respiratory phases were evaluated as separate static anatomies rather than as a continuous delivery–breathing interplay. The framework assumed a perfect two-dimensional target mask, used one beam, ignored setup and range uncertainty, and was not tested with real-time imaging or a physical delivery system.
CLINICAL TAKEAWAY
Reinforcement learning may eventually support phase-aware proton spot delivery that avoids the normal-tissue cost of a full ITV. Current results demonstrate algorithmic potential only; robustness, imaging latency, segmentation error, interplay, machine constraints, and multi-field delivery remain unresolved.