arXiv:2604.04815cs.CLcs.AI2026-04ACL被引 2

动态时间基准测试让大模型在信息变化中验证假新闻,更贴近真实场景。

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection

论文配图:LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection
图 1 · 摘自论文原文
  • 构建随时间更新的动态证据集,模拟真实信息不确定环境。
  • 22个大模型测试显示开源专家混合模型表现媲美甚至超越闭源顶尖系统。
  • 发现模型推理能力存在显著差距,能识别早期不可验证声明的才具真正推理力。

大语言模型(LLMs)的发展使假新闻检测从简单分类演变为复杂推理任务,但现有评估框架未能同步跟进。当前基准静态化,易受基准数据污染(BDC)影响,无法有效评估模型在时间不确定性下的推理能力。为此,我们提出LiveFact——一个持续更新的动态时间感知基准,模拟现实世界中信息不透明的“迷雾”状态。LiveFact采用动态、时间化的证据集,评估模型在信息不断演变且不完整时的推理能力,而非依赖记忆知识。我们设计双模式评估:分类模式用于最终验证,推理模式用于基于证据的推断,并引入专门机制监控BDC。对22个LLM的测试表明,开源的专家混合模型(如Qwen3-235B-A22B)已达到或超越专有顶级系统性能。更重要的是,分析揭示了显著的‘推理差距’:具备强推理能力的模型能在早期数据片段中识别出不可验证的声明,这一特质是传统静态基准所忽略的。LiveFact为评估鲁棒、具有时间感知能力的AI验证设定了可持续标准。

原文摘要 · Abstract (English)

The rapid development of Large Language Models (LLMs) has transformed fake news detection and fact-checking tasks from simple classification to complex reasoning. However, evaluation frameworks have not kept pace. Current benchmarks are static, making them vulnerable to benchmark data contamination (BDC) and ineffective at assessing reasoning under temporal uncertainty. To address this, we introduce LiveFact a continuously updated benchmark that simulates the real-world "fog of war" in misinformation detection. LiveFact uses dynamic, temporal evidence sets to evaluate models on their ability to reason with evolving, incomplete information rather than on memorized knowledge. We propose a dual-mode evaluation: Classification Mode for final verification and Inference Mode for evidence-based reasoning, along with a component to monitor BDC explicitly. Tests with 22 LLMs show that open-source Mixture-of-Experts models, such as Qwen3-235B-A22B, now match or outperform proprietary state-of-the-art systems. More importantly, our analysis finds a significant "reasoning gap." Capable models exhibit epistemic humility by recognizing unverifiable claims in early data slices-an aspect traditional static benchmarks overlook. LiveFact sets a sustainable standard for evaluating robust, temporally aware AI verification.

假新闻检测动态评估大模型推理时间感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。