构建真实场景下的事实核查基准,评估大模型在谣言识别中的表现。
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
- 设计涵盖多模态、跨领域的6000条真实命题,模拟真实误信息场景。
- 引入未知率(UnR)指标,量化模型对不确定性的判断能力。
- 公开测试平台,适合研究事实核查与模型可信度的学者使用。
大型语言模型(LLMs)在推理、证据检索和解释生成方面具备推动事实核查的潜力。然而,现有基准无法全面评估 LLMs 和多模态大语言模型(MLLMs)在真实误信息场景中的表现。为此,我们提出 RealFactBench,一个综合性基准,用于评估 LLMs 与 MLLMs 在知识验证、谣言检测和事件核实等多样化现实任务中的事实核查能力。RealFactBench 包含 6000 条来自权威来源的高质量声明,涵盖多模态内容与多样领域。评估框架引入未知率(UnR)指标,实现对模型处理不确定性能力的更精细评估,平衡过度保守与过度自信。在 7 个代表性 LLMs 与 4 个 MLLMs 上的广泛实验揭示了其在真实场景下的局限性,并为后续研究提供了重要洞见。RealFactBench 已公开于 https://github.com/kalendsyang/RealFactBench.git。
原文摘要 · Abstract (English)
Large Language Models (LLMs) hold significant potential for advancing fact-checking by leveraging their capabilities in reasoning, evidence retrieval, and explanation generation. However, existing benchmarks fail to comprehensively evaluate LLMs and Multimodal Large Language Models (MLLMs) in realistic misinformation scenarios. To bridge this gap, we introduce RealFactBench, a comprehensive benchmark designed to assess the fact-checking capabilities of LLMs and MLLMs across diverse real-world tasks, including Knowledge Validation, Rumor Detection, and Event Verification. RealFactBench consists of 6K high-quality claims drawn from authoritative sources, encompassing multimodal content and diverse domains. Our evaluation framework further introduces the Unknown Rate (UnR) metric, enabling a more nuanced assessment of models' ability to handle uncertainty and balance between over-conservatism and over-confidence. Extensive experiments on 7 representative LLMs and 4 MLLMs reveal their limitations in real-world fact-checking and offer valuable insights for further research. RealFactBench is publicly available at https://github.com/kalendsyang/RealFactBench.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。