用智能体+强化学习提升视频假信息检测的可靠性
FactGuard: Agentic Video Misinformation Detection via Reinforcement Learning
- 让模型像侦探一样迭代验证,主动调用外部工具找证据
- 在三个数据集上均达领先效果,对碎片化证据更敏感
- 适合需要高可信度检测的平台或内容审核场景
多模态大语言模型(MLLMs)虽通过统一多模态推理显著推进了视频假信息检测,但常依赖固定深度推理,过度信任内部假设,尤其在关键证据稀缺、零散或需外部验证时表现不佳。为此,我们提出FactGuard,一种基于MLLMs的智能体框架,将验证过程建模为迭代推理流程。FactGuard显式评估任务模糊性,选择性调用外部工具获取关键证据,实现推理路径的逐步优化。为进一步增强能力,我们设计两阶段训练策略:结合领域特定的智能体监督微调与决策感知强化学习,以优化工具使用并校准风险敏感决策。在FakeSV、FakeTT和FakeVV上的大量实验表明,FactGuard性能达到当前最优,并验证其出色的鲁棒性和泛化能力。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have substantially advanced video misinformation detection through unified multimodal reasoning, but they often rely on fixed-depth inference and place excessive trust in internally generated assumptions, particularly in scenarios where critical evidence is sparse, fragmented, or requires external verification. To address these limitations, we propose FactGuard, an agentic framework for video misinformation detection that formulates verification as an iterative reasoning process built upon MLLMs. FactGuard explicitly assesses task ambiguity and selectively invokes external tools to acquire critical evidence, enabling progressive refinement of reasoning trajectories. To further strengthen this capability, we introduce a two-stage training strategy that combines domain-specific agentic supervised fine-tuning with decision-aware reinforcement learning to optimize tool usage and calibrate risk-sensitive decision making. Extensive experiments on FakeSV, FakeTT, and FakeVV demonstrate FactGuard's state-of-the-art performance and validate its excellent robustness and generalization capacity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。