arXiv:2603.23519cs.CLcs.AI2026-03被引 6

测试大模型在医疗长对话中的记忆与理解能力,发现普遍表现不佳。

MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?

  • 构建了400个真实医疗场景的多轮对话测试集,每轮平均22次
  • 17个前沿模型平均准确率低于60%,最佳仅59.75%
  • 专为评估医疗AI长期记忆和抗干扰能力设计,适合研究者使用

大型语言模型在多个专业领域展现出强大能力,并已应用于医疗等高风险场景。然而,现有医疗基准极少测试实际应用所需的长上下文记忆、干扰鲁棒性及安全防御能力。为此,我们提出了MedMT-Bench,一个模拟完整诊疗过程的挑战性医疗多轮指令遵循基准。通过逐场景数据合成并经专家人工精修,构建出400个高度贴近真实场景的测试用例,每例平均包含22轮对话(最多52轮),涵盖5类复杂指令遵循问题。评估采用基于大模型判卷的协议,结合实例级评分标准与原子测试点,经专家标注验证,人-模型一致性达91.94%。测试17个前沿模型,整体准确率均低于60.00%,最优模型仅为59.75%。MedMT-Bench可成为推动更安全可靠医疗AI研究的关键工具。基准数据已在https://openreview.net/attachment?id=aKyBCsPOHB&name=supplementary_material公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive capabilities across various specialist domains and have been integrated into high-stakes areas such as medicine. However, as existing medical-related benchmarks rarely stress-test the long-context memory, interference robustness, and safety defense required in practice. To bridge this gap, we introduce MedMT-Bench, a challenging medical multi-turn instruction following benchmark that simulates the entire diagnosis and treatment process. We construct the benchmark via scene-by-scene data synthesis refined by manual expert editing, yielding 400 test cases that are highly consistent with real-world application scenarios. Each test case has an average of 22 rounds (maximum of 52 rounds), covering 5 types of difficult instruction following issues. For evaluation, we propose an LLM-as-judge protocol with instance-level rubrics and atomic test points, validated against expert annotations with a human-LLM agreement of 91.94\%. We test 17 frontier models, all of which underperform on MedMT-Bench (overall accuracy below 60.00\%), with the best model reaching 59.75\%. MedMT-Bench can be an essential tool for driving future research towards safer and more reliable medical AI. The benchmark is available in https://openreview.net/attachment?id=aKyBCsPOHB&name=supplementary_material

医疗AI长对话大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。