用大模型自动设计医疗强化学习奖励函数,提升治疗方案优化效果。
medR: Reward Engineering for Clinical Offline Reinforcement Learning via Tri-Drive Potential Functions
- 构建三驱动奖励函数:生存、信心、能力,融合临床知识
- 通过量化指标筛选最优奖励结构,提升策略性能
- 适用于复杂疾病治疗方案优化,适合临床RL研究者
强化学习(RL)为优化动态治疗方案(DTRs)提供了强大框架。然而,临床强化学习的核心瓶颈在于奖励工程:在复杂且稀疏的离线环境中,如何定义安全有效的信号以引导策略学习。现有方法多依赖人工启发式规则,难以跨病种泛化。为此,我们提出一种自动化流程,利用大语言模型(LLMs)进行离线奖励设计与验证。采用由生存、信心和能力三个核心组件构成的势函数形式化奖励函数,并引入定量指标,在部署前严格评估与选择最优奖励结构。通过整合LLM驱动的领域知识,本框架可自动为特定疾病设计奖励函数,显著提升所得策略的性能。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) offers a powerful framework for optimizing dynamic treatment regimes (DTRs). However, clinical RL is fundamentally bottlenecked by reward engineering: the challenge of defining signals that safely and effectively guide policy learning in complex, sparse offline environments. Existing approaches often rely on manual heuristics that fail to generalize across diverse pathologies. To address this, we propose an automated pipeline leveraging Large Language Models (LLMs) for offline reward design and verification. We formulate the reward function using potential functions consisted of three core components: survival, confidence, and competence. We further introduce quantitative metrics to rigorously evaluate and select the optimal reward structure prior to deployment. By integrating LLM-driven domain knowledge, our framework automates the design of reward functions for specific diseases while significantly enhancing the performance of the resulting policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。