构建多模态伪造检测与解释的统一基准,支持跨模态可解释分析。
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark
- 提出基于思维链的三阶段训练法,融合人类对齐评估与推理。
- 涵盖图像、视频、音频等四类内容,支持时空定位与伪造类型追踪。
- 提供空间/时间重叠率等可量化解释指标,适合安全敏感场景研究者。
随着生成式AI快速发展,图像、视频和音频等领域的合成内容日益逼真,加剧了虚假信息传播风险。现有检测方法多集中于二分类,缺乏对伪造行为的详细可解释性说明,限制了其在安全关键场景中的应用。此外,当前方法通常独立处理各模态,缺乏统一的跨模态伪造检测与解释基准。为此,我们提出METER,一个统一的多模态可解释伪造检测基准,覆盖图像、视频、音频及音视频内容。数据集包含四个赛道,不仅要求真假分类,还需提供基于证据链的解释,包括时空定位、文本推理和伪造类型追溯。相比以往基准,METER具备更广的模态覆盖和更丰富的可解释性度量,如空间/时间交并比(IoU)、多类别追踪与证据一致性。我们进一步提出一种人类对齐的三阶段思维链(CoT)训练策略,结合SFT、DPO及新型GRPO阶段,集成人类对齐评估器与推理机制。期望METER能成为生成媒体时代通用可解释伪造检测的标准基础。
原文摘要 · Abstract (English)
With the rapid advancement of generative AI, synthetic content across images, videos, and audio has become increasingly realistic, amplifying the risk of misinformation. Existing detection approaches predominantly focus on binary classification while lacking detailed and interpretable explanations of forgeries, which limits their applicability in safety-critical scenarios. Moreover, current methods often treat each modality separately, without a unified benchmark for cross-modal forgery detection and interpretation. To address these challenges, we introduce METER, a unified, multi-modal benchmark for interpretable forgery detection spanning images, videos, audio, and audio-visual content. Our dataset comprises four tracks, each requiring not only real-vs-fake classification but also evidence-chain-based explanations, including spatio-temporal localization, textual rationales, and forgery type tracing. Compared to prior benchmarks, METER offers broader modality coverage and richer interpretability metrics such as spatial/temporal IoU, multi-class tracing, and evidence consistency. We further propose a human-aligned, three-stage Chain-of-Thought (CoT) training strategy combining SFT, DPO, and a novel GRPO stage that integrates a human-aligned evaluator with CoT reasoning. We hope METER will serve as a standardized foundation for advancing generalizable and interpretable forgery detection in the era of generative media.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。