首个音视频联合编辑框架,解决生成内容不同步与语义冲突问题
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing

- 采用双向跨模态交互机制,实现音视频同步对齐
- 在基准测试中同步指标提升30%,语义一致性提升12.5%
- 适合需要高质量音视频协同生成的创作者与研究人员
物理世界中的视觉与听觉事件天然耦合,但现有视频编辑方法多采用解耦流程,缺乏双向模态交互,导致两大缺陷:(i) 音视频不同步,(ii) 生成音频与保留内容存在语义冲突。为此,我们提出SpongeBob,首个端到端音视频联合编辑框架,支持双向跨模态交互。为保证同步性,设计同步感知机制,通过双向注意力、时间对齐与空间约束对齐视觉编辑与声学事件;为保障上下文一致性,引入上下文感知模块,利用声学与视觉上下文注意力防止语义冲突。此外,提出同步保持训练与引导(SPTG)以增强对齐效果而不降低质量。由于配对数据稀缺,构建可扩展的数据流水线及大规模主体级数据集,并提出SpongeBob-Bench用于系统评估。实验表明,SpongeBob显著优于现有基线,同步指标Sync-C提升30%,语义一致性指标Ctx-F1提升12.5%。
原文摘要 · Abstract (English)
Visual and acoustic events in the physical world are inherently coupled, yet existing video editing methods typically adopt decoupled pipelines, lacking bidirectional modality interaction. This results in two key limitations: (i) audio-visual desynchronization and (ii) contextual conflicts between generated audio and preserved content. To address these, we propose SpongeBob, the first end-to-end audio-visual joint editing framework featuring bidirectional cross-modal interaction. For synchronization, a Sync-Aware Mechanism aligns visual edits with sound events via bidirectional attention, temporal alignment, and spatial constraints. For contextual consistency, a Context-Aware Module leverages acoustic and visual context attention to prevent semantic clashes. Additionally, we introduce Sync-Preserving Training and Guidance (SPTG) to enhance alignment without degrading quality. Due to the scarcity of paired data, we construct a scalable data pipeline and a large-scale subject-level dataset. We also propose SpongeBob-Bench for systematic evaluation. Experiments show SpongeBob significantly outperforms existing baselines, improving Sync-C by 30% and Ctx-F1 by 12.5%. Our project page is available at: https://hy-spongebob.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。