arXiv:2605.11814cs.AI2026-05被引 1

构建医疗记忆评估基准,揭示主流模型在长期诊疗中的记忆衰退问题。

MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare

论文配图:MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
图 1 · 摘自论文原文
  • 基于临床患者原型生成真实医疗轨迹,构建2000会话、1.6万轮交互数据集。
  • 提出动态流式评估协议,模拟生产环境中的持续记忆积累过程。
  • 发现复杂医学推理与噪声鲁棒性是主流模型的核心短板,适合医疗AI研发者参考。

个性化健康代理的大规模部署需要精确、安全且具备长期临床追踪能力的记忆机制。然而现有基准主要关注日常开放域对话,难以反映真实医疗应用的高风险复杂性。基于服务千万级用户的行业领先健康管理代理的严苛生产需求,我们提出MedMemoryBench。通过人机协同流程,基于临床基础的合成患者原型生成高度真实的长时程医疗轨迹,构建包含约2000个会话和16000次交互轮次的专家验证数据集。关键在于,MedMemoryBench突破传统静态评估,首创“边构建边评估”的流式评估协议,精准模拟生产环境中动态记忆累积过程。此外,我们系统化研究了记忆饱和现象——持续信息输入主动削弱检索与推理鲁棒性。全面评测揭示主流架构在复杂医学推理与噪声鲁棒性方面存在严重瓶颈。该基准为开发可投入生产的稳健医疗代理奠定了重要基础。

原文摘要 · Abstract (English)

The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, existing benchmarks primarily focus on daily open-domain conversations, failing to capture the high-stakes complexity of real-world medical applications. Motivated by the stringent production requirements of an industry-leading health management agent serving tens of millions of active users, we introduce MedMemoryBench. We develop a human-agent collaborative pipeline to synthesize highly realistic, long-horizon medical trajectories based on clinically grounded, synthetic patient archetypes. This process yields a massive, expertly validated dataset comprising approximately 2,000 sessions and 16,000 interaction turns. Crucially, MedMemoryBench departs from traditional static evaluations by pioneering an "evaluate-while-constructing" streaming assessment protocol, which precisely mirrors dynamic memory accumulation in production environments. Furthermore, we formalize and systematically investigate the critical phenomenon of memory saturation, where sustained information influx actively degrades retrieval and reasoning robustness. Comprehensive benchmarking reveals severe bottlenecks in mainstream architectures, particularly concerning complex medical reasoning and noise resilience. By exposing these fundamental flaws, MedMemoryBench establishes a vital foundation for developing robust, production-ready medical agents.

医疗AI记忆机制评估基准长时追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。