arXiv:2607.25543cs.CVcs.AI2026-07

分离音视频检测,更准识别通用场景的AI伪造内容。

Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

论文配图:Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
图 1 · 摘自论文原文
  • 音视频独立分析,避免跨模态假设带来的误判
  • 多粒度视觉+双分支音频结构,精准捕捉伪造痕迹
  • 在国际竞赛中夺冠,适合做通用AI造假检测

生成式AI已将音视频伪造从人脸深度伪造扩展到通用场景。现有方法依赖音视频内容一致性来识别伪造,但我们在通用场景中发现这一假设并不总是成立。因此我们主张,决策级融合比特征级融合更稳健。为此提出DAV-Det,一种解耦的音视频AIGC检测系统,分别独立建模各模态的取证证据。视觉检测器利用全局、块级和片段级的多粒度表征,捕捉空间伪造线索;音频检测器则通过门控时序-频谱双分支结构,建模时间与频谱不规则性。本方法在IJCAI-ECAI 2026 DDL 2.0 Workshop的通用AIGC音视频检测挑战赛中排名第一,最终得分为0.8460。代码已开源:https://github.com/tuffy-studio/DAV-Det。

原文摘要 · Abstract (English)

Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes. Existing AIGC detection methods assume audio-visual content correspondence, identifying forgeries by spotting cross-modal inconsistencies. However, we empirically find that this assumption does not consistently hold in general scenarios. We argue that, for general audio-visual AIGC detection, decision-level fusion is a more robust alternative to feature-level fusion. Therefore, we propose DAV-Det, a decoupled audio-visual AIGC detection system that independently models forensic evidence from each modality. The visual detector leverages multi-granularity representations at global, patch, and segment levels to capture spatial forgery cues, while the audio detector exploits both temporal and spectral irregularities via a gated temporal-spectral dual-branch architecture to model acoustic artifacts. Our method ranks 1st in the General AIGC Audio-Video Detection Challenge of the IJCAI-ECAI 2026 DDL 2.0 Workshop, with a final score of 0.8460. Code is available at https://github.com/tuffy-studio/DAV-Det.

AI伪造检测音视频融合多模态分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。