arXiv:2605.19766cs.CLcs.AI2026-05中稿 · AAMAS 2026被引 1

用大模型生成长期医疗对话数据,评估医生智能体的记忆能力。

Synthesis and Evaluation of Long-term History-aware Medical Dialogue

  • 分三阶段合成患者病史与多轮对话,构建真实长时医疗对话数据集。
  • 现有顶尖大模型在跨会话推理任务中表现仍差,暴露记忆短板。
  • 适合研究医疗AI记忆、长期推理与对话系统评测的团队使用。

有效的医疗智能体必须能够回顾并推理患者的长期医疗历史。然而,缺乏具有真实长期对话时间线的数据集,限制了系统性评估。真实临床文本受隐私与伦理约束,而现有基准集中于孤立交互,无法捕捉跨会话推理。我们提出一种基于大模型的高质长期医疗对话合成框架。方法包括:基于知识引导的三阶段设计——构建具有多样化疾病与并发症轨迹的合成患者档案,每就诊生成多轮对话,最终整合为连贯的长期历史数据集MediLongChat。我们设立三项基准任务:会话内推理、跨会话推理与合成推理,以评估医疗智能体的记忆能力。为评估数据质量,引入多维度评价框架,结合向量度量与大模型作为裁判的评估。具体定义自动指标:忠实性、连贯性与多样性,并引入两项基于大模型的评估:正确性与真实性。基准实验显示,即使最先进的大模型在MediLongChat上也表现不佳。这些发现验证了该基准的适用性,凸显了发展针对性方法以推进医疗智能体的必要性。

原文摘要 · Abstract (English)

An effective healthcare agent must be able to recall and reason over a patient's longitudinal medical history. However, the absence of datasets with realistic long-term dialogue timelines limits systematic evaluation. Real clinical text is constrained by privacy and ethics, while existing benchmarks focus on isolated interactions, failing to capture cross-session reasoning. We introduce a framework for synthesizing high-quality, long-term medical dialogues with LLMs. Our approach entails a knowledge-guided decomposition into three stages: constructing synthetic patient profiles with diverse disease and complication trajectories, generating multi-turn dialogues per encounter, and integrating them into a coherent longitudinal history dataset, MediLongChat. We establish three benchmark tasks-In-dialogue Reasoning, Cross-dialogue Reasoning, and Synthesis Reasoning-to evaluate the memory capabilities of healthcare agents. To assess data quality, we introduce a multi-dimensional evaluation framework combining vector-based metrics with LLM-as-a-judge assessments. Specifically, we define automatic measures-Faithfulness, Coherence, and Diversity-together with two LLM-based evaluations: Correctness and Realism. Benchmark experiments show that even state-of-the-art LLMs struggle with MediLongChat. These findings highlight the benchmark's applicability and underscore the need for tailored methods to advance healthcare agents.

医疗对话长时记忆大模型评估数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。