用深度推理检测视频假信息,让AI像人一样思考。
Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

- 结合规则强化学习与多阶段训练,让模型学会推理假信息
- 在超10万对视频文本数据上训练,实现可解释的判断
- 适合需要可信检测系统的安全、媒体和平台方
社交媒体上多模态虚假信息迅速传播,引发广泛关注,但视频虚假信息检测研究受限于缺乏大规模、多样化的数据集。现有方法常过度依赖固定模板,缺乏对欺骗性内容的深层推理能力。为此,我们提出 FakeVV,一个包含超过10万对视频-文本的数据集,具备细粒度、可解释的标注。同时,我们进一步提出 Fact-R1 框架,融合深度推理与协作式规则强化学习。Fact-R1 通过三阶段训练:(1)虚假信息长链式思维(CoT)指令微调,(2)基于直接偏好优化(DPO)的偏好对齐,(3)使用新型可验证奖励函数的群体相对策略优化(GRPO)。该方法使 Fact-R1 在复杂多模态虚假信息场景中展现出类似先进文本强化学习系统中的涌现推理行为。本工作建立了一种新范式,连接大规模视频理解、推理引导对齐与可解释验证。
原文摘要 · Abstract (English)
The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive content. To address these challenges, we introduce FakeVV, a large-scale benchmark comprising over 100,000 video-text pairs with fine-grained, interpretable annotations. In addition, we further propose Fact-R1, a novel framework that integrates deep reasoning with collaborative rule-based reinforcement learning. Fact-R1 is trained through a three-stage process: (1) misinformation long-Chain-of-Thought (CoT) instruction tuning, (2) preference alignment via Direct Preference Optimization (DPO), and (3) Group Relative Policy Optimization (GRPO) using a novel verifiable reward function. This enables Fact-R1 to exhibit emergent reasoning behaviors comparable to those observed in advanced text-based reinforcement learning systems, but in the more complex multimodal misinformation setting. Our work establishes a new paradigm for misinformation detection, bridging large-scale video understanding, reasoning-guided alignment, and interpretable verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。