用真实医生提问对比大模型,发现好答案更重清晰与细节。
MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences
- 让医生直接比较模型回答,基于真实临床问题做选择
- 三款模型领跑,偏好侧重内容深度与表达清晰度
- 适合关注临床实用性、想验证模型真实表现的研究者
大型语言模型(LLMs)正日益融入临床决策支持、医学教育和患者沟通等医疗工作流。然而,现有评估方法多依赖静态模板化基准,难以反映真实临床实践的复杂性与动态性,导致评测结果与实际临床价值脱节。为此,我们提出MedArena——一个交互式评估平台,允许医生使用自身真实的医疗问题测试并比较主流LLMs。在12个模型中收集了1571条偏好数据(截至2025年11月1日),结果显示Gemini 2.0 Flash Thinking、Gemini 2.5 Pro和GPT-4o位列前三。仅约三分之一问题为事实记忆类(如MedQA),多数涉及治疗选择、临床记录或患者沟通,约20%为多轮对话。医生在解释偏好时更看重内容深度与呈现清晰度,而非单纯准确率。控制响应长度与格式后,模型排名仍稳定,证明结果可靠性。通过以真实临床需求为基础的评估,MedArena为衡量和提升医疗LLMs的实用性提供了可扩展的方案。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly central to clinician workflows, spanning clinical decision support, medical education, and patient communication. However, current evaluation methods for medical LLMs rely heavily on static, templated benchmarks that fail to capture the complexity and dynamics of real-world clinical practice, creating a dissonance between benchmark performance and clinical utility. To address these limitations, we present MedArena, an interactive evaluation platform that enables clinicians to directly test and compare leading LLMs using their own medical queries. Given a clinician-provided query, MedArena presents responses from two randomly selected models and asks the user to select the preferred response. Out of 1571 preferences collected across 12 LLMs up to November 1, 2025, Gemini 2.0 Flash Thinking, Gemini 2.5 Pro, and GPT-4o were the top three models by Bradley-Terry rating. Only one-third of clinician-submitted questions resembled factual recall tasks (e.g., MedQA), whereas the majority addressed topics such as treatment selection, clinical documentation, or patient communication, with ~20% involving multi-turn conversations. Additionally, clinicians cited depth and detail and clarity of presentation more often than raw factual accuracy when explaining their preferences, highlighting the importance of readability and clinical nuance. We also confirm that the model rankings remain stable even after controlling for style-related factors like response length and formatting. By grounding evaluation in real-world clinical questions and preferences, MedArena offers a scalable platform for measuring and improving the utility and efficacy of medical LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。