arXiv:2506.13474cs.CLcs.AI2025-06被引 12

用强化学习训练语言代理,模拟医生循证诊断过程。

Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learning

  • 构建可迭代提问与解读检验的假设驱动语言代理
  • 在MIMIC-CDM上诊断准确率提升,测试次数减少23%
  • 适合想提升智能问诊系统交互能力的研究者

临床决策是一个动态、交互且循环的过程,医生需反复决定执行何种临床操作,并根据新发现的信息进行诊断与治疗。大型语言模型(LLMs)有潜力支持这一过程,但现有应用大多存在两种局限:要么假设所有患者信息即时可用,未建模交互式、迭代的探查过程;要么仅依赖预训练模型的“开箱即用”能力,缺乏任务定制化训练。与此不同,我们提出一种假设驱动的、不确定性感知的语言代理LA-CDM,通过反复请求并解释相关检测来逐步收敛至诊断。采用监督与强化学习结合的混合训练范式,对LA-CDM进行训练,以实现三大目标:准确生成假设、估计假设不确定性、高效决策。我们在涵盖四种腹腔疾病的实世界数据集MIMIC-CDM上评估方法,证明了显式训练临床决策能显著提升诊断性能与效率。

原文摘要 · Abstract (English)

Clinical decision-making is a dynamic, interactive, and cyclic process where doctors have to repeatedly decide on which clinical action to perform and consider newly uncovered information for diagnosis and treatment. Large Language Models (LLMs) have the potential to support clinicians in this process, however, most applications of LLMs in clinical decision support suffer from one of two limitations: Either they assume the unrealistic scenario of immediate availability of all patient information and do not model the interactive and iterative investigation process, or they restrict themselves to the limited "out-of-the-box" capabilities of large pre-trained models without performing task-specific training. In contrast to this, we propose to model clinical decision-making for diagnosis with a hypothesis-driven uncertainty-aware language agent, LA-CDM, that converges towards a diagnosis via repeatedly requesting and interpreting relevant tests. Using a hybrid training paradigm combining supervised and reinforcement learning, we train LA-CDM with three objectives targeting critical aspects of clinical decision-making: accurate hypothesis generation, hypothesis uncertainty estimation, and efficient decision-making. We evaluate our methodology on MIMIC-CDM, a real-world dataset covering four abdominal diseases containing various clinical tests and show the benefit of explicitly training clinical decision-making for increasing diagnostic performance and efficiency.

临床决策强化学习语言代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。