arXiv:2607.16916cs.LG2026-07

用强化学习动态优化膀胱癌复发治疗,让方案随病情变化自适应调整。

Enhancing Personalized Bladder Cancer Treatment Through Reinforcement Learning: A Recurrent Patient State Transition Decision Support Framework

  • 构建基于MDP与DQN的患者状态转移模型,逐阶段优化治疗决策。
  • 在模拟环境中实现63,918.87累计奖励和6.62%策略提升,效果优于现有方法。
  • 生成可解释的治疗路径日志,适合临床医生与AI辅助肿瘤学研究者使用。

膀胱癌治疗需个性化且动态调整,尤其在复发情况下,治疗效果随临床阶段变化。传统决策支持系统依赖静态指南或单步预测模型,难以捕捉疾病进展。本文提出一种基于强化学习的复发性患者状态转移模拟框架,整合预测性状态转移建模、马尔可夫决策过程(MDP)与深度Q网络(DQN),通过模拟患者轨迹实现治疗序列优化。预测模块评估治疗后肿瘤特征变化,强化学习代理则依据动态临床状态持续优化决策。该框架支持动态、个体化治疗规划,并生成可解释的治疗路径与详细仿真日志,增强临床透明度。在对比实验中,该框架获得63,918.87的累积奖励、每轮平均训练损失0.0056,以及6.62%的策略改进率,验证了其在模拟复发治疗环境中的有效性和鲁棒性。结果表明,该方法为个性化膀胱癌治疗与人工智能辅助精准肿瘤学提供了灵活的决策支持框架。

原文摘要 · Abstract (English)

Bladder cancer treatment requires personalized and adaptive decision-making, particularly for recurrent disease, where treatment effectiveness changes across successive clinical episodes. Conventional clinical decision support systems typically rely on static treatment guidelines or single-step predictive models, limiting their ability to capture disease progression over time. This paper presents a recurrent patient state-transition simulation framework for bladder cancer treatment planning that integrates predictive state-transition modeling with a Markov Decision Process (MDP) and a Deep Q-Network (DQN) reinforcement learning environment. The predictive module estimates changes in tumor characteristics following treatment, while the reinforcement learning agent sequentially optimizes treatment decisions by interacting with simulated patient trajectories. This framework enables dynamic, patient-specific treatment planning by continuously adapting recommendations to evolving clinical states. It also generates interpretable treatment trajectories and detailed simulation logs to improve transparency and support clinical decision-making. The proposed framework was evaluated against existing reinforcement learning-based treatment planning approaches. It achieved a cumulative reward of 63,918.87, an average training loss per episode of 0.0056, and a policy improvement score of 6.62%, demonstrating effective sequential learning and robust treatment optimization in a simulated recurrent treatment environment. These findings highlight the potential of recurrent patient state-transition simulation with reinforcement learning as a flexible decision-support framework for personalized bladder cancer treatment planning and AI-assisted precision oncology.

强化学习精准医疗肿瘤治疗决策支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。