用迭代树状推理法,精准揪出医学长文本里的虚假论断
Iterative Tree Analysis for Medical Critics
- 构建迭代树状推理框架,分步拆解并验证复杂医学主张
- 在医学事实核查任务中准确率比前人方法提升10%
- 适合医疗AI安全、可信生成研究者使用
大型语言模型在医学领域应用广泛,但易产生幻觉。开放域长篇医学文本中的误导性批评声明难以验证,原因有二:一是关键主张常隐含于文本深层,仅靠表层提取无效;二是基于标记的检索常缺乏精确证据,无法支撑验证。本文提出一种名为迭代树分析(Iterative Tree Analysis, ITA)的新方法,可从长医学文本中提取隐含主张,并通过迭代自适应的树状推理过程进行精准验证。该过程结合自上而下的任务分解与自下而上的证据整合,实现机制层面的深度推理。大量实验表明,ITA在复杂医学文本的事实核查任务中显著优于先前方法,准确率提升10%。此外,我们将公开一个全面的测试集,以推动该领域的研究进展。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been widely adopted across various domains, yet their application in the medical field poses unique challenges, particularly concerning the generation of hallucinations. Hallucinations in open-ended long medical text manifest as misleading critical claims, which are difficult to verify due to two reasons. First, critical claims are often deeply entangled within the text and cannot be extracted based solely on surface-level presentation. Second, verifying these claims is challenging because surface-level token-based retrieval often lacks precise or specific evidence, leaving the claims unverifiable without deeper mechanism-based analysis. In this paper, we introduce a novel method termed Iterative Tree Analysis (ITA) for medical critics. ITA is designed to extract implicit claims from long medical texts and verify each claim through an iterative and adaptive tree-like reasoning process. This process involves a combination of top-down task decomposition and bottom-up evidence consolidation, enabling precise verification of complex medical claims through detailed mechanism-level reasoning. Our extensive experiments demonstrate that ITA significantly outperforms previous methods in detecting factual inaccuracies in complex medical text verification tasks by 10%. Additionally, we will release a comprehensive test set to the public, aiming to foster further advancements in research within this domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。