arXiv:2607.19201cs.CLcs.AI2026-07

构建首个临床诊断推理基准,评估模型是否真能基于正确证据做判断。

MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams

  • 基于西班牙住院医师考试案例,标注证据片段与论点关系
  • 三阶段评测:找证据、抽论点、判支持/攻击关系,覆盖1024个病例
  • 首次提供巴斯克语版本,助力多语言医学AI发展

临床自然语言处理评估仍以多选题为主,仅衡量最终答案正确性,无法判断模型是否基于相关、缺失或矛盾的证据得出正确诊断。我们提出MIRA-Ev,一个基于西班牙医学住院医师考试(MIR)案例的临床论证挖掘基准,由专家医生重新标注了句子级前提、主张及有向支持/攻击关系,并同步发布西班牙语(母语)、英语和巴斯克语版本,是首个巴斯克语临床论证资源。MIRA-Ev采用三级任务体系:证据句检索、论证成分抽取与关系分类,推动更细致的临床推理评估。

原文摘要 · Abstract (English)

Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot detect when a model reaches the correct diagnosis while grounding it in irrelevant, absent, or contradictory evidence. We introduce MIRA-Ev, a clinical argument mining benchmark built on Spanish Médico Interno Residente (MIR) licensing-exam cases, re-annotated by expert clinicians with span-level premises, claims, and directed support/attack relations, and released in parallel Spanish (native), English, and Basque versions, the first clinical argumentation resource in Basque. MIRA-Ev organizes evaluation into a three-tier task hierarchy: evidence sentence retrieval, argumentative component extraction, and relation classification.

临床NLP论证挖掘多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。