arXiv:2604.23957cs.CV2026-04中稿 · the 34th ACM Inter…被引 1

LAVA通过融合音视频水印,实现压缩和错位下的鲁棒深伪检测与定位

LAVA: Layered Audio-Visual Anti-tampering Watermarking for Robust Deepfake Detection and Localization

  • 音视频水印分层融合并动态对齐,抵抗压缩与异步干扰
  • 在压缩和错位下仍保持99.9%的检测准确率(AP=0.999)
  • 适合用于短视频平台的深伪内容溯源与防篡改系统

主动水印为短视频中的深伪篡改检测与定位提供了有前景的解决方案。然而,现有方法常将音视频证据分离,并假设水印信号在真实世界退化下依然可靠,导致篡改定位易受多模态错位和压缩失真影响。此外,现有半脆弱视觉水印方法在编码压缩下性能显著下降,因其嵌入频带与压缩敏感区域重叠。为此,我们提出分层音视频抗篡改水印框架LAVA,通过跨模态水印融合与校准感知对齐,在压缩和音视频不同步条件下保持一致可靠的篡改证据,实现鲁棒的篡改定位。大量实验表明,LAVA达到近乎完美的检测性能(AP=0.999),对压缩和多模态错位具有强鲁棒性,且显著优于现有音视频融合基线的篡改定位可靠性。

原文摘要 · Abstract (English)

Proactive watermarking offers a promising approach for deepfake tamper detection and localization in short-form videos. However, existing methods often decouple audio and visual evidence and assume that watermark signals remain reliable under real-world degradations, making tamper localization vulnerable to multimodal misalignment and compression distortions. Moreover, existing semi-fragile visual watermarking methods often degrade significantly under codec compression because their embedding bands overlap with compression-sensitive frequency regions. To address these limitations, we propose Layered Audio-Visual Anti-tampering Watermarking (LAVA), a calibration-aware audio-visual watermark fusion framework for deepfake tamper detection and localization. LAVA leverages cross-modal watermark fusion and calibration-aware alignment to preserve consistent and reliable tamper evidence under compression and audio-visual asynchrony, enabling robust tamper localization. Extensive experiments demonstrate that LAVA achieves near-perfect detection performance (AP = 0.999), remains robust to compression and multimodal misalignment, and significantly improves tamper localization reliability over existing audio-visual fusion baselines.

深伪检测音视频融合水印技术鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。