构建多语言临床文本基准,评估大模型真实医疗场景表现。
BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text
- 覆盖9种语言、87项任务,涵盖诊疗全流程
- 开源模型性能可比商用模型,旧医模不如新通用模型
- 适合医疗AI研发者与评测人员参考
大语言模型在医疗应用中潜力巨大,但现有评估多基于考试题或科研文献,难以反映电子健康记录等真实临床数据的复杂性。为填补这一空白,我们推出BRIDGE——一个涵盖九种语言、87项任务的综合性多语言基准,覆盖六类临床阶段和20个应用场景,包括分诊、会诊、信息抽取、诊断、预后和编码等,涉及14个专科领域。我们系统评估了95个大模型(含DeepSeek-R1、GPT-4o、Gemini系列、Qwen3系列)在不同推理策略下的表现。结果显示,模型规模、语言、任务类型和专科差异显著影响性能。值得注意的是,开源模型表现可达商用模型水平,而基于旧架构的医学微调模型常落后于新通用模型。BRIDGE及其排行榜为真实临床文本理解的模型开发与评估提供基础参考。排行榜链接:https://huggingface.co/spaces/YLab-Open/BRIDGE-Medical-Leaderboard
原文摘要 · Abstract (English)
Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on large-scale real-world data such as electronic health records (EHRs) is critical, as clinical decisions are directly informed by these sources, yet current evaluations remain limited. Most existing benchmarks rely on medical exam-style questions or PubMed-derived text, failing to capture the complexity of real-world clinical data. Others focus narrowly on specific application scenarios, limiting their generalizability across broader clinical use. To address this gap, we present BRIDGE, a comprehensive multilingual benchmark comprising 87 tasks sourced from real-world clinical data sources across nine languages. It covers eight major task types spanning the entire continuum of patient care across six clinical stages and 20 representative applications, including triage and referral, consultation, information extraction, diagnosis, prognosis, and billing coding, and involves 14 clinical specialties. We systematically evaluated 95 LLMs (including DeepSeek-R1, GPT-4o, Gemini series, and Qwen3 series) under various inference strategies. Our results reveal substantial performance variation across model sizes, languages, natural language processing tasks, and clinical specialties. Notably, we demonstrate that open-source LLMs can achieve performance comparable to proprietary models, while medically fine-tuned LLMs based on older architectures often underperform versus updated general-purpose models. The BRIDGE and its corresponding leaderboard serve as a foundational resource and a unique reference for the development and evaluation of new LLMs in real-world clinical text understanding. The BRIDGE leaderboard: https://huggingface.co/spaces/YLab-Open/BRIDGE-Medical-Leaderboard
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。