arXiv:2512.20822cs.CLcs.AI2025-12ACL被引 3

构建医疗大模型评估新基准,发现并修复幻觉与错误反转问题

MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs

  • 基于MIMIC-IV与UMLS构建真实患者场景下的医学知识库
  • 发现当前模型普遍存在幻觉支持和真理颠倒等致命缺陷
  • 提出CoRFu方法,提升16.4点准确率并消除安全错误

大型语言模型在医疗领域应用日益广泛,但其可靠性和安全性仍存疑。现有评估或孤立测试医学知识,或仅评估患者层面推理而无法验证正确性,存在明显空白。我们提出MediEval,将MIMIC-IV电子健康记录与基于UMLS等生物医学词典构建的统一知识库关联,生成多样化的事实与反事实医学陈述,实现对知识溯源与上下文一致性的四象限系统评估。通过该框架,我们识别出当前专有、开源及领域专用模型普遍存在的幻觉支持与真理颠倒等关键失败模式。为此,我们提出反事实风险感知微调(CoRFu),一种基于DPO的不对称惩罚方法,专门针对高风险混淆进行优化。CoRFu相比基线模型提升16.4点宏F1分数,并彻底消除真理颠倒错误,显著提高准确率与安全性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly applied to medicine, yet their adoption is limited by concerns over reliability and safety. Existing evaluations either test factual medical knowledge in isolation or assess patient-level reasoning without verifying correctness, leaving a critical gap. We introduce MediEval, a benchmark that links MIMIC-IV electronic health records (EHRs) to a unified knowledge base built from UMLS and other biomedical vocabularies. MediEval generates diverse factual and counterfactual medical statements within real patient contexts, enabling systematic evaluation across a 4-quadrant framework that jointly considers knowledge grounding and contextual consistency. Using this framework, we identify critical failure modes, including hallucinated support and truth inversion, that current proprietary, open-source, and domain-specific LLMs frequently exhibit. To address these risks, we propose Counterfactual Risk-Aware Fine-tuning (CoRFu), a DPO-based method with an asymmetric penalty targeting unsafe confusions. CoRFu improves by +16.4 macro-F1 points over the base model and eliminates truth inversion errors, demonstrating both higher accuracy and substantially greater safety.

医疗LLM知识推理模型安全评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。