arXiv:2602.10687cs.CVcs.AI2026-02中稿 · ICML被引 2

统一检测图文视频伪造并定位,解决多模态任务中简单任务主导难题。

OmniVL-Guard: Towards Unified Vision-Language Forgery Detection and Grounding via Balanced RL

  • 用自进化思维链生成高质量推理路径,克服冷启动问题。
  • 动态调节奖励尺度和任务权重,实现检测与定位的平衡优化。
  • 在跨域场景下零样本泛化能力强,适合真实世界反虚假信息应用。

现有伪造检测方法多局限于单模态或双模态设置,难以应对现实世界中交错出现的文本、图像和视频。为此,本文提出统一框架OmniVL-Guard,实现对多模态伪造内容的联合检测与定位。在该统一设置下,不同模态间的交互以及同时进行检测与定位的双重需求导致显著的“难度偏差”问题:简单的真伪分类任务会主导梯度更新,损害细粒度定位性能。为此,提出基于平衡强化学习的OmniVL-Guard框架,包含两项核心设计:自进化思维链生成(Self-Evolving CoT Generation)与自适应奖励缩放策略优化(Adaptive Reward Scaling Policy Optimization, ARSPO)。前者生成高质量推理路径,缓解冷启动挑战;后者动态调整奖励尺度与任务权重,确保多任务联合优化均衡。大量实验表明,OmniVL-Guard显著优于现有最优方法,并在跨领域场景中展现出零样本鲁棒泛化能力。数据集与代码已公开于https://github.com/shen8424/OmniVL-Guard。

原文摘要 · Abstract (English)

Existing forgery detection methods are often limited to uni-modal or bi-modal settings, failing to handle the interleaved text, images, and videos prevalent in real-world misinformation. To bridge this gap, this paper targets to develop a unified framework for omnibus vision-language forgery detection and grounding. In this unified setting, the {interplay} between diverse modalities and the dual requirements of simultaneous detection and localization pose a critical ``difficulty bias`` problem: the simpler veracity classification task tends to dominate the gradients, leading to suboptimal performance in fine-grained grounding during multi-task optimization. To address this challenge, we propose \textbf{OmniVL-Guard}, a balanced reinforcement learning framework for omnibus vision-language forgery detection and grounding. Particularly, OmniVL-Guard comprises two core designs: Self-Evolving CoT Generatio and Adaptive Reward Scaling Policy Optimization (ARSPO). {Self-Evolving CoT Generation} synthesizes high-quality reasoning paths, effectively overcoming the cold-start challenge. Building upon this, {Adaptive Reward Scaling Policy Optimization (ARSPO)} dynamically modulates reward scales and task weights, ensuring a balanced joint optimization. Extensive experiments demonstrate that OmniVL-Guard significantly outperforms state-of-the-art methods and exhibits zero-shot robust generalization across out-of-domain scenarios. The dataset and code are publicly available at https://github.com/shen8424/OmniVL-Guard.

伪造检测多模态强化学习视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。