arXiv:2512.00048cs.ROcs.AI2025-12中稿 · AAAI

用因果强化学习提升机器人照护痴呆患者的决策能力

Causal Reinforcement Learning based Agent-Patient Interaction with Clinical Domain Knowledge

  • 引入因果图建模患者状态与行动的动态关系
  • 在模拟场景中实现更高奖励与更稳定患者状态
  • 无需微调大模型,轻量部署即可生成临床对齐对话

强化学习在自适应医疗干预(如痴呆照护)中面临数据稀缺、决策需可解释、患者状态动态复杂且具因果性等挑战。本文提出因果结构感知强化学习(CRL),将因果发现与推理显式融入策略优化。该方法使智能体学习并利用有向无环图(DAG)描述人类行为状态与机器人动作间的因果依赖,从而实现更高效、可解释且鲁棒的决策。我们在模拟机器人辅助认知照护场景中验证方法,智能体与具有动态情绪、认知及参与度状态的虚拟患者交互。实验表明,CRL智能体相比传统无模型强化学习基线,在累积奖励、维持理想患者状态一致性方面表现更优,且行为可解释、符合临床逻辑。其性能优势在不同权重策略与超参数设置下依然稳健。此外,我们展示了一种轻量级LLM部署方案:将固定策略嵌入系统提示,基于推断状态映射到行动,无需微调即可生成一致、支持性的对话。本工作展示了因果强化学习在人机交互应用中的潜力,尤其适用于对可解释性、自适应性和数据效率要求高的场景。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) faces significant challenges in adaptive healthcare interventions, such as dementia care, where data is scarce, decisions require interpretability, and underlying patient-state dynamic are complex and causal in nature. In this work, we present a novel framework called Causal structure-aware Reinforcement Learning (CRL) that explicitly integrates causal discovery and reasoning into policy optimization. This method enables an agent to learn and exploit a directed acyclic graph (DAG) that describes the causal dependencies between human behavioral states and robot actions, facilitating more efficient, interpretable, and robust decision-making. We validate our approach in a simulated robot-assisted cognitive care scenario, where the agent interacts with a virtual patient exhibiting dynamic emotional, cognitive, and engagement states. The experimental results show that CRL agents outperform conventional model-free RL baselines by achieving higher cumulative rewards, maintaining desirable patient states more consistently, and exhibiting interpretable, clinically-aligned behavior. We further demonstrate that CRL's performance advantage remains robust across different weighting strategies and hyperparameter settings. In addition, we demonstrate a lightweight LLM-based deployment: a fixed policy is embedded into a system prompt that maps inferred states to actions, producing consistent, supportive dialogue without LLM finetuning. Our work illustrates the promise of causal reinforcement learning for human-robot interaction applications, where interpretability, adaptiveness, and data efficiency are paramount.

强化学习因果推理医疗机器人大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。