arXiv:2507.14632cs.CV2025-07被引 26

用强化学习统一检测图文生成内容并解释原因,效果超越传统方法。

BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM

  • 采用单阶段强化学习,不依赖监督微调,提升模型探索能力。
  • 在图像和视频检测上均达到当前最佳性能,跨模态能力显著增强。
  • 适合关注生成内容检测与可解释性的研究人员和安全团队。

生成式AI的快速发展显著提升了图像与视频合成质量,加剧了多模态视觉误导风险。尽管近期多模态大模型(MLLM)在透明化检测生成内容方面展现出推理与解释潜力,但现有方法大多将图像与视频取证视为孤立任务,未充分挖掘跨模态协同效应。为此,我们提出 extbf{BusterX++},一个统一的MLLM,实现图像与视频的联合检测与可解释推理。同时引入 extbf{GenBuster-Bench++},一个精心构建、难度对齐的基准,包含均衡的图像与视频样本,覆盖近期生成模型及多样真实场景。在该受控环境下,我们重新审视广泛采用的 $SFT \rightarrow RL$ 后训练范式。结果表明,仅使用稀疏奖励信号的单阶段纯强化学习策略,在统一与单模态设置下均表现持平或超越强基线。关键洞察是:监督微调(SFT)降低策略熵,限制搜索空间,抑制探索;而单阶段纯强化学习维持更高策略熵,有效激发图像与视频取证间的自发跨模态能力迁移。大量实验验证了BusterX++的领先性能,凸显强化学习在统一多模态视觉推理中的强大潜力。

原文摘要 · Abstract (English)

The rapid advancement of generative AI has substantially improved image and video synthesis, amplifying the risk of multimodal visual misinformation. Recent MLLMs have shown promise for transparent AI-generated content detection through reasoning and explanation, yet existing approaches largely treat image and video forensics as isolated tasks, leaving cross-modal synergies underexplored. To address this, we present \textbf{BusterX++}, a unified MLLM for joint image and video detection with interpretable reasoning. We also introduce \textbf{GenBuster-Bench++}, a meticulously curated, difficulty-aligned benchmark containing balanced image and video samples spanning recent generation models and diverse real-world scenarios. Using this controlled setting, we revisit the widely adopted $SFT \rightarrow RL$ post-training paradigm. Notably, our findings demonstrate that a single-stage, pure RL strategy driven strictly by sparse outcome rewards consistently matches or surpasses a strong SFT+RL baseline across both unified and single-modality settings. Our key insight reveals that SFT imposes lower policy entropy, which restricts the policy search space and dampens exploratory freedom. In contrast, single-stage pure RL maintains higher policy entropy throughout training, effectively unlocking the spontaneous emergence of cross-modal capability transfer between image and video forensics. Extensive experiments demonstrate that BusterX++ achieves state-of-the-art performance, highlighting the powerful potential of RL for unified cross-modal visual reasoning.

生成内容检测多模态强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。