用AI聊天机器人帮病人理解医疗信息,无需人工标注数据。
Chatbot To Help Patients Understand Their Health
- 基于多智能体LLM与强化学习,通过模拟出院场景训练。
- 在模拟测试中表现优于非专业人类,对话清晰有结构。
- 适合医疗健康领域,可低成本复用于其他开放对话场景。
患者需要具备足够知识以主动参与自身护理。我们提出NoteAid-Chatbot,一种基于多智能体大语言模型和强化学习(RL)的对话式AI,采用创新的‘对话即学习’框架,在无真人标注数据条件下提升患者理解力。该系统基于轻量级LLaMA 3.2 3B模型,分两阶段训练:首先在合成生成的医学对话数据上进行监督微调,再通过奖励信号驱动强化学习,奖励来自模拟医院出院场景中的患者理解评估。评估包含全面的人类对齐测试与案例研究,结果显示NoteAid-Chatbot展现出清晰性、相关性和结构化对话等关键涌现行为,尽管未显式训练这些属性。结果表明,即使使用简单的近端策略优化(PPO)奖励建模,也能有效训练出轻量、领域专用的多轮对话机器人,支持多样化教育策略并达成复杂沟通目标。图灵测试显示其表现超越非专家人类。尽管当前聚焦医疗,本框架证明了低成本PPO强化学习在真实、开放对话领域的可行性与潜力,拓展了基于强化学习对齐方法的应用边界。
原文摘要 · Abstract (English)
Patients must possess the knowledge necessary to actively participate in their care. We present NoteAid-Chatbot, a conversational AI that promotes patient understanding via a novel 'learning as conversation' framework, built on a multi-agent large language model (LLM) and reinforcement learning (RL) setup without human-labeled data. NoteAid-Chatbot was built on a lightweight LLaMA 3.2 3B model trained in two stages: initial supervised fine-tuning on conversational data synthetically generated using medical conversation strategies, followed by RL with rewards derived from patient understanding assessments in simulated hospital discharge scenarios. Our evaluation, which includes comprehensive human-aligned assessments and case studies, demonstrates that NoteAid-Chatbot exhibits key emergent behaviors critical for patient education, such as clarity, relevance, and structured dialogue, even though it received no explicit supervision for these attributes. Our results show that even simple Proximal Policy Optimization (PPO)-based reward modeling can successfully train lightweight, domain-specific chatbots to handle multi-turn interactions, incorporate diverse educational strategies, and meet nuanced communication objectives. Our Turing test demonstrates that NoteAid-Chatbot surpasses non-expert human. Although our current focus is on healthcare, the framework we present illustrates the feasibility and promise of applying low-cost, PPO-based RL to realistic, open-ended conversational domains, broadening the applicability of RL-based alignment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。