arXiv:2412.18947cs.CLcs.AI2024-12AAAI被引 24

构建医疗大模型幻觉评估新基准,提升AI诊疗可靠性。

MedHallBench: A New Benchmark for Assessing Hallucination in Medical Large Language Models

  • 结合专家验证病例与医学数据库,构建严谨评估数据集。
  • 采用自动化评分与临床专家双轨验证,发现传统指标不足。
  • 适合医疗AI研发者、评测人员及临床安全研究人员使用。

医疗大语言模型(MLLMs)在医疗应用中展现出潜力,但其易产生医学上不合理的虚构信息,对患者安全构成重大风险。本文提出MedHallBench,一个全面的基准框架,用于评估和缓解MLLMs的幻觉问题。该框架整合专家验证的医疗案例与权威医学数据库,构建高可信度评估数据集;采用自动化的医学影像描述幻觉测量(ACHMI)评分系统,并结合临床专家严格评估,利用强化学习实现自动标注。通过专为医疗场景优化的基于人类反馈的强化学习(RLHF)训练流程,实现在多样临床情境下的全面模型评估,同时确保高标准准确性。对比实验涵盖多个模型,建立了主流大语言模型的基线表现。结果表明,相较于传统指标,ACHMI能更细致地揭示幻觉影响,凸显其在幻觉评估中的优势。本研究为提升医疗大模型可靠性提供了基础框架,并提出了应对医疗AI幻觉的关键策略。

原文摘要 · Abstract (English)

Medical Large Language Models (MLLMs) have demonstrated potential in healthcare applications, yet their propensity for hallucinations -- generating medically implausible or inaccurate information -- presents substantial risks to patient care. This paper introduces MedHallBench, a comprehensive benchmark framework for evaluating and mitigating hallucinations in MLLMs. Our methodology integrates expert-validated medical case scenarios with established medical databases to create a robust evaluation dataset. The framework employs a sophisticated measurement system that combines automated ACHMI (Automatic Caption Hallucination Measurement in Medical Imaging) scoring with rigorous clinical expert evaluations and utilizes reinforcement learning methods to achieve automatic annotation. Through an optimized reinforcement learning from human feedback (RLHF) training pipeline specifically designed for medical applications, MedHallBench enables thorough evaluation of MLLMs across diverse clinical contexts while maintaining stringent accuracy standards. We conducted comparative experiments involving various models, utilizing the benchmark to establish a baseline for widely adopted large language models (LLMs). Our findings indicate that ACHMI provides a more nuanced understanding of the effects of hallucinations compared to traditional metrics, thereby highlighting its advantages in hallucination assessment. This research establishes a foundational framework for enhancing MLLMs' reliability in healthcare settings and presents actionable strategies for addressing the critical challenge of AI hallucinations in medical applications.

医疗AI幻觉检测评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。