arXiv:2505.24830cs.CLcs.AI2025-05被引 4

通过原子事实核查提升医疗问答的准确性和可解释性。

Improving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs

  • 将生成答案拆分为可验证的原子事实,逐个比对医学指南库。
  • 事实准确率提升40%,幻觉检测率达50%。
  • 适合需要高可信度医疗AI的临床场景与监管审查。

大型语言模型(LLMs)虽具备广泛医学知识,但易产生幻觉和错误引用,制约其临床应用与合规性。现有检索增强生成(RAG)方法虽部分缓解此问题,但幻觉与低粒度可解释性仍存。本文提出一种新型原子事实核查框架,将LLM生成的回答分解为离散、可验证的原子事实,每个事实独立比对权威医学指南知识库。该方法支持精准纠错并可追溯至原始文献片段,显著提升医疗问答的事实准确率与可解释性。在多位医学专家参与的多阅读评估及自动化开放问答基准测试中,框架实现最高40%的整体答案改进,幻觉检测率达50%。每条原子事实均可回溯至数据库中最相关的文本块,提供细粒度透明解释,填补当前医疗AI应用的关键空白。本工作推动了更可信、可靠的临床级LLM应用,满足临床落地关键前提,增强对AI辅助医疗的信任。

原文摘要 · Abstract (English)

Large language models (LLMs) exhibit extensive medical knowledge but are prone to hallucinations and inaccurate citations, which pose a challenge to their clinical adoption and regulatory compliance. Current methods, such as Retrieval Augmented Generation, partially address these issues by grounding answers in source documents, but hallucinations and low fact-level explainability persist. In this work, we introduce a novel atomic fact-checking framework designed to enhance the reliability and explainability of LLMs used in medical long-form question answering. This method decomposes LLM-generated responses into discrete, verifiable units called atomic facts, each of which is independently verified against an authoritative knowledge base of medical guidelines. This approach enables targeted correction of errors and direct tracing to source literature, thereby improving the factual accuracy and explainability of medical Q&A. Extensive evaluation using multi-reader assessments by medical experts and an automated open Q&A benchmark demonstrated significant improvements in factual accuracy and explainability. Our framework achieved up to a 40% overall answer improvement and a 50% hallucination detection rate. The ability to trace each atomic fact back to the most relevant chunks from the database provides a granular, transparent explanation of the generated responses, addressing a major gap in current medical AI applications. This work represents a crucial step towards more trustworthy and reliable clinical applications of LLMs, addressing key prerequisites for clinical application and fostering greater confidence in AI-assisted healthcare.

医疗问答事实核查可解释性LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。