arXiv:2601.04160cs.CLcs.CE2026-01ACL被引 12

构建金融虚假信息检测新基准,揭示大模型在无参考时的推理缺陷。

All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection

  • 基于段落级新闻设计无参考检测与对比诊断双任务
  • 有对比时模型准确率显著提升,无参考时预测不稳且错误频发
  • 适合研究模型可信推理与金融信息真实性验证的学者

我们提出RFC Bench,一个用于评估大语言模型在真实新闻场景下金融虚假信息检测能力的基准。该基准以段落为单位,捕捉金融新闻中依赖分散线索生成语义的上下文复杂性。基准定义了两个互补任务:无参考虚假信息检测与基于成对原始/扰动输入的对比诊断。实验表明,当提供对比上下文时性能明显增强;而在无参考设置下,模型暴露显著弱点,包括预测不稳定和无效输出增多。这些结果表明,当前模型在缺乏外部依据时难以维持连贯信念状态。通过揭示这一差距,RFC Bench为研究无参考推理与推进真实场景下更可靠的金融虚假信息检测提供了结构化测试平台。

原文摘要 · Abstract (English)

We introduce RFC Bench, a benchmark for evaluating large language models on financial misinformation under realistic news. RFC Bench operates at the paragraph level and captures the contextual complexity of financial news where meaning emerges from dispersed cues. The benchmark defines two complementary tasks: reference free misinformation detection and comparison based diagnosis using paired original perturbed inputs. Experiments reveal a consistent pattern: performance is substantially stronger when comparative context is available, while reference free settings expose significant weaknesses, including unstable predictions and elevated invalid outputs. These results indicate that current models struggle to maintain coherent belief states without external grounding. By highlighting this gap, RFC Bench provides a structured testbed for studying reference free reasoning and advancing more reliable financial misinformation detection in real world settings.

虚假信息检测大模型评估金融AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。