用知识图谱自动检测医疗大模型回答的真假,更准更可信。
Assessing Automated Fact-Checking for Medical LLM Responses with Knowledge Graphs
- 将回答拆成原子事实,通过医学知识图谱验证证据链。
- 与医生判断相关性显著提升,能区分不同模型能力差异。
- 结果可解释,适合医疗AI安全评估与模型改进使用。
大型语言模型(LLMs)在医疗任务中展现强大能力,但其在高风险医疗场景中的部署需严格验证。本文研究基于医学知识图谱(KG)的自动化事实性评估方法可行性。为此,提出FAITH框架,无需参考答案即可将响应分解为原子陈述,链接至医学知识图谱,并依据证据路径进行评分。在多种医疗任务上的实验表明,该方法与临床医生判断具有显著更高的相关性,能有效区分不同能力水平的LLMs,且对文本变化具有鲁棒性。其评分过程具备内在可解释性,有助于理解并缓解当前大模型的局限。结论表明,尽管存在局限,利用知识图谱是医疗领域自动化事实评估的重要方向。
原文摘要 · Abstract (English)
The recent proliferation of large language models (LLMs) holds the potential to revolutionize healthcare, with strong capabilities in diverse medical tasks. Yet, deploying LLMs in high-stakes healthcare settings requires rigorous verification and validation to understand any potential harm. This paper investigates the reliability and viability of using medical knowledge graphs (KGs) for the automated factuality evaluation of LLM-generated responses. To ground this investigation, we introduce FAITH, a framework designed to systematically probe the strengths and limitations of this KG-based approach. FAITH operates without reference answers by decomposing responses into atomic claims, linking them to a medical KG, and scoring them based on evidence paths. Experiments on diverse medical tasks with human subjective evaluations demonstrate that KG-grounded evaluation achieves considerably higher correlations with clinician judgments and can effectively distinguish LLMs with varying capabilities. It is also robust to textual variances. The inherent explainability of its scoring can further help users understand and mitigate the limitations of current LLMs. We conclude that while limitations exist, leveraging KGs is a prominent direction for automated factuality assessment in healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。