评测大模型在临床问答中的表现,覆盖七种题型。
Overview of the ClinIQLink 2025 Shared Task on Medical Question-Answering
- 构建7类医学问答数据集,覆盖真假、选择、多跳等题型。
- 使用精确匹配和嵌入评分双指标评估模型,得分超90%的交由医生审核。
- 适合关注医疗大模型评估与可信度的研究者参考。
本文介绍了ClinIQLink 2025共享任务,该任务与ACL 2025第24届BioNLP研讨会同期举行,旨在对大语言模型(LLMs)在面向全科医生水平的医学问答能力进行压力测试。任务提供4,978个专家验证、基于医学来源的问题-答案对,涵盖七种格式:是非题、单选题、无序列表、简答、反向简答、多跳推理及反向多跳推理。参赛系统以Docker或Apptainer镜像形式提交,在CodaBench平台或马里兰大学Zaratan集群上运行。任务1通过精确匹配评估封闭式问题,开放性问题采用三级嵌入评分机制;任务2由医师小组对表现最优的模型输出进行人工审计。
原文摘要 · Abstract (English)
In this paper, we present an overview of ClinIQLink, a shared task, collocated with the 24th BioNLP workshop at ACL 2025, designed to stress-test large language models (LLMs) on medically-oriented question answering aimed at the level of a General Practitioner. The challenge supplies 4,978 expert-verified, medical source-grounded question-answer pairs that cover seven formats: true/false, multiple choice, unordered list, short answer, short-inverse, multi-hop, and multi-hop-inverse. Participating systems, bundled in Docker or Apptainer images, are executed on the CodaBench platform or the University of Maryland's Zaratan cluster. An automated harness (Task 1) scores closed-ended items by exact match and open-ended items with a three-tier embedding metric. A subsequent physician panel (Task 2) audits the top model responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。