为合成音视频伪造内容嵌入跨模态水印,实现真实音频恢复与篡改定位
Cross-Modal Watermarking for Authentic Audio Recovery and Tamper Localization in Synthesized Audiovisual Forgeries
- 在视觉中嵌入真实音频水印,实现伪造内容溯源
- 在多种语音克隆和唇同步攻击下仍能准确恢复原始音频
- 适合反虚假信息、数字取证等安全场景
语音克隆与唇同步模型的发展催生了合成音视频伪造(SAVFs),使虚假内容更具欺骗性。现有方法虽能检测或定位篡改,却无法恢复原始音频语义。本文提出真实音频恢复(AAR)与音轨篡改定位(TLA)任务,并设计跨模态水印框架,在生成前将真实音频嵌入视觉内容。该方法在多种语音克隆与唇同步攻击下均表现优异,可有效抵御音视频虚假信息传播。
原文摘要 · Abstract (English)
Recent advances in voice cloning and lip synchronization models have enabled Synthesized Audiovisual Forgeries (SAVFs), where both audio and visuals are manipulated to mimic a target speaker. This significantly increases the risk of misinformation by making fake content seem real. To address this issue, existing methods detect or localize manipulations but cannot recover the authentic audio that conveys the semantic content of the message. This limitation reduces their effectiveness in combating audiovisual misinformation. In this work, we introduce the task of Authentic Audio Recovery (AAR) and Tamper Localization in Audio (TLA) from SAVFs and propose a cross-modal watermarking framework to embed authentic audio into visuals before manipulation. This enables AAR, TLA, and a robust defense against misinformation. Extensive experiments demonstrate the strong performance of our method in AAR and TLA against various manipulations, including voice cloning and lip synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。