A tandem reinforcement learning framework for localized prostate cancer treatment planning and machine parameter optimization.

Shaffer N; Mudireddy AR; St-Aubin J

doi:10.1002/mp.70306

← 뒤로

A tandem reinforcement learning framework for localized prostate cancer treatment planning and machine parameter optimization.

1/5 보강

Medical physics 📖 저널 OA 33.8% 2022~2026 2026 Vol.53(2) p. e70306

PICO 자동 추출 (휴리스틱, conf 2/4)

유사 논문

P · Population 대상 환자/모집단

20 patients and compared to reference plans optimized with a commercial TPS.

I · Intervention 중재 / 시술

추출되지 않음

C · Comparison 대조 / 비교

추출되지 않음

O · Outcome 결과 / 결론

The algorithm rapidly generates VMAT prostate cancer treatment plans that meet clinical constraints and are dosimetrically comparable to manually optimized plans without the use of a commercial TPS optimizer. This work demonstrates the feasibility of RL as a tool to fully automate the VMAT planning process, offering the potential to decrease planning times while maintaining plan quality.

Shaffer N, Mudireddy AR, St-Aubin J

📖 무료 전문 🟢 PMC 전문 PMC12857532

PubMed ↗ DOI ↗ BibTeX ↓ RIS ↓

📝 환자 설명용 한 줄

이 논문을 인용하기

↓ .bib ↓ .ris

APA Shaffer N, Mudireddy AR, St-Aubin J (2026). A tandem reinforcement learning framework for localized prostate cancer treatment planning and machine parameter optimization.. Medical physics, 53(2), e70306. https://doi.org/10.1002/mp.70306

MLA Shaffer N, et al.. "A tandem reinforcement learning framework for localized prostate cancer treatment planning and machine parameter optimization.." Medical physics, vol. 53, no. 2, 2026, pp. e70306.

PMID 41615038 ↗

DOI 10.1002/mp.70306

Abstract

[BACKGROUND] Volumetric modulated arc therapy (VMAT) machine parameter optimization (MPO) is a complex, high-dimensional problem typically solved with inverse planning solutions that are both temporally and computationally expensive. While machine learning techniques have been explored to automate this process, they often supplement rather than replace conventional optimizers and are fundamentally limited by the quality and diversity of training data. Reinforcement learning (RL) offers a promising alternative, finding optimal strategies through trial-and-error by maximizing a narrowly tailored reward function, which can potentially discover novel solutions beyond mimicking features present in existing plans.

[PURPOSE] The purpose of this study was to develop and validate a deep reinforcement learning-based VMAT MPO algorithm capable of automatically generating clinically comparable treatment plans for prostate cancer that meet machine constraints, entirely independent of a commercial treatment planning system (TPS) optimizer.

[METHODS] A dataset comprised of 100 prostate cancer patients planned using the criteria from PACE-B SBRT arm serve as the basis for network training using a 70-10-20 training/validation/testing split. An RL framework using a Proximal Policy Optimization (PPO) algorithm was developed to train two tandem convolutional neural networks that sequentially optimize multi-leaf collimator (MLC) positions and monitor units (MUs) using current dose, contoured structure masks, and current machine parameters as inputs. Training was designed to predict MLC positions and MUs that maximize a dose-volume histogram (DVH)-based reward function tailored to prioritize meeting clinical objectives. The fully trained networks were executed on a test set of 20 patients and compared to reference plans optimized with a commercial TPS.

[RESULTS] The RL algorithm generated plans in an average of 6.3 ± 4.7 s. Compared to the reference plans, the RL-generated plans demonstrated improved sparing for both the bladder and rectum across their respective dosimetric endpoints. When normalizing to 95% coverage, the RL generated plans resulted in a statistically significant increase in the PTV , while achieving a significantly reduced for the rectum. All RL plans successfully satisfied all clinical objectives used to optimize the reference plans.

[CONCLUSIONS] We successfully developed and validated a deep RL framework for VMAT MPO. The algorithm rapidly generates VMAT prostate cancer treatment plans that meet clinical constraints and are dosimetrically comparable to manually optimized plans without the use of a commercial TPS optimizer. This work demonstrates the feasibility of RL as a tool to fully automate the VMAT planning process, offering the potential to decrease planning times while maintaining plan quality.

🏷️ 키워드 / MeSH 📖 같은 키워드 OA만

🏷️ 같은 키워드 · 무료전문 — 이 논문 MeSH/keyword 기반

A Phase I Study of Hydroxychloroquine and Suba-Itraconazole in Men with Biochemical Relapse of Prostate Cancer (HITMAN-PC): Dose Escalation Results.
Cancer research communications 2026 Talmor B 외 📖 OA
Self-management of male urinary symptoms: qualitative findings from a primary care trial.
The British journal of general practice : the journal of the Royal College of General Practitioners 2026 Wheeler JR 외 📖 OA
Clinical and Liquid Biomarkers of 20-Year Prostate Cancer Risk in Men Aged 45 to 70 Years.
JAMA network open 2026 Lindholz M 외 📖 OA
Diagnostic accuracy of Ga-PSMA PET/CT versus multiparametric MRI for preoperative pelvic invasion in the patients with prostate cancer.
Science progress 2026 Qin Z 외 📖 OA
Association of patient health education with the postoperative health related quality of life in low- intermediate recurrence risk differentiated thyroid cancer patients.
Scientific reports 2026 Li S 외 📖 OA
Early local immune activation following intra-operative radiotherapy in human breast tissue.
Oncoimmunology 2026 Tiefenthaller A 외 📖 OA

이 논문을 인용하기

Abstract 한글 요약

🏷️ 키워드 / MeSH 📖 같은 키워드 OA만

🏷️ 같은 키워드 · 무료전문 — 이 논문 MeSH/keyword 기반

Abstract