arXiv:2510.09275cs.CLcs.AI2025-10ACL被引 4

提出动态医学诊断评估框架,更真实检验大模型在复杂临床场景下的表现。

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation

  • 构建动态病例生成机制,模拟真实问诊中的混淆因素。
  • 测试发现顶尖大模型在复杂场景下准确率大幅下降,暴露严重弱点。
  • 不仅看准确率,还评估真实性、帮助性和一致性,更贴近临床需求。

医学诊断是高风险且复杂的领域,对患者治疗至关重要。然而,当前大语言模型(LLMs)的评估仍局限于基于公开考试题目的基准,存在数据污染偏差,且忽略真实问诊中超越教科书案例的混杂因素。虽然动态评估提供了新方向,但多数仍不足以支撑以诊断为导向的评测,覆盖范围有限,且仅关注准确率。为此,我们提出 DyReMe,一个动态医学诊断基准,可进行可控且可扩展的诊断鲁棒性压力测试。与静态考试题不同,DyReMe 生成全新、咨询风格的病例,引入临床相关的混淆因素,如鉴别诊断和常见误诊因素,并变化表达风格以捕捉患者描述的多样性。除准确率外,还评估模型在真实性、有用性和一致性三个临床相关维度的表现。实验表明,该动态方法带来更具挑战性的评估,揭示了先进 LLM 在临床混杂场景下的显著缺陷。研究强调,亟需基于临床实际混淆因素的评估框架来衡量可信医学诊断能力。

原文摘要 · Abstract (English)

Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLMs) remain limited in capturing key challenges of clinical diagnostic scenarios. Most rely on benchmarks derived from public exams, raising contamination bias that can inflate performance, and they overlook the confounded nature of real consultations beyond textbook cases. Recent dynamic evaluations offer a promising alternative, but often remain insufficient for diagnosis-oriented benchmarking, with limited coverage of clinically grounded confounders and trustworthiness beyond accuracy. To address these gaps, we propose DyReMe, a dynamic benchmark for medical diagnostics that provides a controlled and scalable stress test of diagnostic robustness. Unlike static exam-style questions, DyReMe generates fresh, consultation-style cases that incorporate clinically grounded confounders, such as differential diagnoses and common misdiagnosis factors. It also varies expression styles to capture heterogeneous patient-style descriptions. Beyond accuracy, DyReMe evaluates LLMs on three additional clinically relevant dimensions: veracity, helpfulness, and consistency. Our experiments show that this dynamic approach yields more challenging assessments and exposes substantial weaknesses of stateof-the-art LLMs under clinically confounded diagnostic settings. These findings highlight the urgent need for evaluation frameworks that better assess trustworthy medical diagnostics 1 under clinically grounded confounders.

医学AI动态评估大模型评测诊断系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。