arXiv:2601.12661cs.AI2026-01被引 2

评测医疗AI咨询的全流程能力,发现高准确率背后的信息收集短板。

MedConsultBench: A Full-Cycle, Fine-Grained, Process-Aware Benchmark for Medical Consultation Agents

  • 用细粒度信息单元追踪每轮对话中的临床信息获取过程。
  • 19个大模型中多数诊断准但信息收集效率低且用药安全差。
  • 适合想提升医疗AI临床实践能力的研究者和开发者。

当前医疗咨询智能体评估多聚焦结果导向任务,常忽视端到端流程完整性和临床安全性。尽管近年出现动态交互基准,但仍碎片化且粗粒度,无法捕捉专业咨询所需的结构化追问逻辑与诊断严谨性。为此,我们提出MedConsultBench,一个覆盖问诊、诊断、治疗规划及随访问答全周期的综合评估框架。通过引入原子信息单元(AIUs),在子对话轮次层面追踪临床信息获取,实现对22项细粒度指标的精确监控。该基准解决在线咨询中的模糊性问题,评估模型在保持简洁的同时具备不确定性感知的追问能力,并强调药物方案兼容性及遵循约束的后续计划修订能力。对19个大语言模型的系统评估显示,高诊断准确率常掩盖信息收集效率低下与用药安全隐患。这一发现揭示了理论医学知识与临床实践能力间的显著差距,确立了MedConsultBench作为推动医疗AI贴近真实临床需求的坚实基础。

原文摘要 · Abstract (English)

Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential for real-world practice. While recent interactive benchmarks have introduced dynamic scenarios, they often remain fragmented and coarse-grained, failing to capture the structured inquiry logic and diagnostic rigor required in professional consultations. To bridge this gap, we propose MedConsultBench, a comprehensive framework designed to evaluate the complete online consultation cycle by covering the entire clinical workflow from history taking and diagnosis to treatment planning and follow-up Q\&A. Our methodology introduces Atomic Information Units (AIUs) to track clinical information acquisition at a sub-turn level, enabling precise monitoring of how key facts are elicited through 22 fine-grained metrics. By addressing the underspecification and ambiguity inherent in online consultations, the benchmark evaluates uncertainty-aware yet concise inquiry while emphasizing medication regimen compatibility and the ability to handle realistic post-prescription follow-up Q\&A via constraint-respecting plan revisions. Systematic evaluation of 19 large language models reveals that high diagnostic accuracy often masks significant deficiencies in information-gathering efficiency and medication safety. These results underscore a critical gap between theoretical medical knowledge and clinical practice ability, establishing MedConsultBench as a rigorous foundation for aligning medical AI with the nuanced requirements of real-world clinical care.

医疗AI评估基准对话系统临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。