arXiv:2503.07937cs.AI2025-03被引 7

用LLM行为分析找出支持与反驳科学观点的证据

LLM-based Corroborating and Refuting Evidence Retrieval for Scientific Claim Verification

  • 通过多角度提问检测LLM响应一致性,识别证据
  • 在多个领域验证中表现优于传统RAG方法
  • 无需模型内部信息,适合各类科学领域使用

本文提出CIBER(基于证据检索的主张调查),扩展了检索增强生成(RAG)框架,用于识别科学主张的支持与反驳证据。CIBER通过分析不同探针下的LLM响应一致性,应对大型语言模型固有的不确定性。该方法不依赖模型内部信息,适用于白盒与黑盒模型,且为无监督学习,可轻松泛化至多种科学领域。在具有不同语言能力的LLM上进行的综合评估显示,CIBER性能显著优于传统RAG方法。研究结果不仅验证了CIBER的有效性,也为未来基于LLM的科学主张验证提供了重要启示。

原文摘要 · Abstract (English)

In this paper, we introduce CIBER (Claim Investigation Based on Evidence Retrieval), an extension of the Retrieval-Augmented Generation (RAG) framework designed to identify corroborating and refuting documents as evidence for scientific claim verification. CIBER addresses the inherent uncertainty in Large Language Models (LLMs) by evaluating response consistency across diverse interrogation probes. By focusing on the behavioral analysis of LLMs without requiring access to their internal information, CIBER is applicable to both white-box and black-box models. Furthermore, CIBER operates in an unsupervised manner, enabling easy generalization across various scientific domains. Comprehensive evaluations conducted using LLMs with varying levels of linguistic proficiency reveal CIBER's superior performance compared to conventional RAG approaches. These findings not only highlight the effectiveness of CIBER but also provide valuable insights for future advancements in LLM-based scientific claim verification.

科学验证LLM应用证据检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。