arXiv:2605.23113cs.CV2026-05中稿 · CVPR

提出IaMSB模型,精准定位音视频伪造区间并抑制跨模态噪声。

Inconsistency-aware Multimodal Schrödinger Bridge for Deepfake Localization

  • 基于薛定谔桥框架,联合估计跨模态一致性与定位区间。
  • 在多个数据集上提升[email protected]达3%~10%,尤其擅长单侧伪造检测。
  • 通过不对称步长分配与瓶颈交互,有效抑制噪声传播。

音视频深度伪造定位需要输出时间区间作为时序证据。尽管已有进展,但单侧或异步伪造场景下的对称融合会引入跨模态噪声,降低高精度定位效果。本文提出IaMSB,一种不一致感知的多模态薛定谔桥(Schrödinger Bridge, SB)方法,可联合估计跨模态一致性并实现区间级定位。与扩散模型不同,薛定谔桥无需显式加噪或去噪,直接最小化路径分布差异并生成一致性分数。IaMSB将一致性估计、跨模态信息选择与桥接步骤调度统一于一个框架:轻量级粗粒度桥首先生成候选区间并估计一致性;基于统计结果选择跨模态见证信号,并不对称分配桥接步数;精炼桥再进行步长调优融合,输出时间对齐的优化区间。该方法能提前识别单侧和异步伪造,通过瓶颈式跨模态交互与步长分配策略,抑制噪声传递、避免冗余迭代。在多个基准测试中,IaMSB稳定提升严格IoU边界精度,使[email protected]提升3%~10%,显著改善高精度定位性能,尤其适用于单侧伪造场景。

原文摘要 · Abstract (English)

Audio-visual deepfake localization demands interval-level outputs that serve as temporal evidence. Despite recent progress, symmetric fusion under single-sided or asynchronous forgeries propagates cross-modal noise, degrading high-precision localization. We present IaMSB, an inconsistency-aware multimodal Schrödinger Bridge (SB) that jointly estimates cross-modal consistency and performs interval-level localization. Unlike diffusion models, SB minimizes path-distribution discrepancy and yields consistency scores without explicit noise injection or denoising. With the Schrödinger Bridge (SB), IaMSB unifies consistency estimation, cross-modal information selection, and bridge-step scheduling in one framework. Specifically, a lightweight coarse bridge first proposes candidate intervals and estimates cross-modal consistency; these statistics select cross-modal witness signals and allocate bridge steps asymmetrically across modalities. A refinement bridge then performs step-tuned fusion and outputs refined, time-aligned intervals. IaMSB anticipates single-sided and asynchronous forgeries and, using bottlenecked cross-modal interaction with step allocation, suppresses noise transfer, avoids unnecessary iterations. Across benchmarks, IaMSB stabilizes strict-IoU boundary precision, raising [email protected] by 3%~10%, and yields improved high-precision localization, particularly for single-sided forgeries.

深度伪造多模态定位薛定谔桥

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。