arXiv:2604.01461cs.AI2026-04被引 1

通过论文间对比发现异常,减少大模型在科研文献分析中的幻觉。

Reducing Hallucinations in LLM-based Scientific Literature Analysis Using Peer Context Outlier Detection

  • 利用同领域论文的相似性检测异常结论
  • 跨领域实验显示98%的异常检测准确率
  • 适合需要高可信度文献分析的研究者

大语言模型在从大量文本中提取数据时易产生幻觉,影响准确性。现有方法如提示工程仅关注单篇文档,忽略文档间的关联。本文提出同伴上下文异常检测(P-COD),通过比较论文间实验设置与结论的相似性,识别不一致信息。若某篇论文的结论与其他同类论文显著不同,则降低其置信度并标记供人工复核;高置信度结果则因获得同行验证而被认为可靠。在6个科学领域的实验表明,该方法在异常检测上达到最高98%的精确率,有效减少幻觉、提升自动化系统的可信度,并帮助研究人员聚焦于模糊案例,优化数据提取流程。

原文摘要 · Abstract (English)

Reducing hallucinations in Large Language Models (LLMs) is essential for improving the accuracy of data extraction from large text corpora. Current methods, like prompt engineering and chain-of-thought prompting, focus on individual documents but fail to consider relationships across a corpus. This paper introduces Peer Context Outlier Detection (P-COD), a novel approach that uses the relationships between documents to improve extraction accuracy. Our application domain is in scientific literature summarization, where papers with similar experiment settings should draw similar conclusions. By comparing extracted data to validated peer information within the corpus, we adjust confidence scores and flag low-confidence results for expert review. High-confidence results, supported by peer validation, are considered reliable. Our experiments demonstrate up to 98% precision in outlier detection across 6 domains of science, demonstrating that our design reduces hallucinations, enhances trust in automated systems, and allows researchers to focus on ambiguous cases, streamlining the data extraction workflows.

大模型幻觉文献分析异常检测科研自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。