arXiv:2605.20197cs.CL2026-05

首个评估医学隐含概念提取的基准,强调推理与证据支撑。

MedicalBench: Evaluating Large Language Models Toward Improved Medical Concept Extraction

论文配图:MedicalBench: Evaluating Large Language Models Toward Improved Medical Concept Extraction
图 1 · 摘自论文原文
  • 将概念提取转为验证任务,结合句子级证据定位。
  • 包含隐含正例、混淆负例,挑战模型深层理解能力。
  • 适合医疗AI研发者,推动可解释性医学语言模型发展。

从电子病历中提取医学概念是诸多下游应用的基础,但因医学概念常隐含于文本中而极具挑战。现有基准多聚焦显式概念,忽略隐含推理。我们提出MedicalBench,一个基于MIMIC-IV出院记录与人工验证的ICD-10编码的基准,通过多阶段大语言模型筛选、医学标注与专家评审构建。该数据集刻意包含隐含正例、语义混淆负例及模型与专家意见不一致的案例。定义两项互补任务:(1)医学概念提取,(2)句子级证据检索,评估结果正确性与可解释性。对主流LLM的评测显示性能仍较低,且与病历长度无关,说明其真正衡量的是推理难度而非表层干扰。MedicalBench首次系统化评估隐含、有证据支撑的医学概念提取,为开发可解释、符合医学事实的语言模型提供基础。

原文摘要 · Abstract (English)

Medical concept extraction from electronic health records underpins many downstream applications, yet remains challenging because medically meaningful concepts are frequently implied rather than explicitly stated in medical narratives. Existing benchmarks with human-annotated evidence spans underscore the importance of grounding extracted concepts in medical text. However, they predominantly focus on explicitly stated concepts instead of implicit concepts. We present MedicalBench, a benchmark for medical concept extraction with evidence grounding that evaluates implicit medical reasoning. MedicalBench formulates medical concept extraction as a verification task over medical note-concept pairs, coupled with sentence-level evidence identification. Built from MIMIC-IV discharge summaries and human-verified ICD-10 codes, the dataset is curated through a multi-stage large language model (LLM) triage pipeline followed by medical annotation and expert review. It deliberately includes implicit positives, semantically confusable negatives, and cases where LLM judgments disagree with medical expert assessments. We define two complementary evaluation tasks: (1) medical concept extraction and (2) sentence-level evidence retrieval, enabling assessment of both correctness and interpretability. Benchmarking state-of-the-art LLMs reveals that performance remains modest, highlighting the difficulty of extracting implicitly expressed concepts. We further show that performance is largely invariant to note length, indicating that MedicalBench isolates reasoning difficulty rather than superficial confounders. MedicalBench provides the first systematic benchmark for implicit, evidence-grounded medical concept extraction, offering a foundation for developing medical language models that can both identify medically relevant concepts and justify their predictions in a transparent and medically faithful manner.

医学AI概念提取大模型评测证据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。