arXiv:2605.18467cs.CV2026-05被引 3

让音视频同步编辑,指令一出自动匹配内容。

InstructAV2AV: Instruction-Guided Audio-Video Joint Editing

论文配图:InstructAV2AV: Instruction-Guided Audio-Video Joint Editing
图 1 · 摘自论文原文
  • 用指令引导音视频联合编辑,保持内容一致。
  • 在11项指标上超越现有方法,数据集达8万对。
  • 适合需要精准音视频控制的创作者使用。

近期基于扩散模型的方法在视频内容操作方面取得了显著进展,但通常忽略伴随音频,导致音视频脱节。本文提出 InstructAV2AV,首个端到端的指令引导音视频联合编辑框架。我们构建了大规模音视频编辑数据集 InsAVE-80K,包含8万对高质量源-目标样本,并设计可扩展的数据合成管道。基于此,我们改进音视频生成主干网络,将音视频输入与噪声潜在码拼接以锚定源上下文,提出源-指令门控注意力机制提升指令遵循与内容保留能力,并采用两阶段训练策略有效迁移预训练先验。大量实验表明,InstructAV2AV 在两个评测集上11项指标均优于现有方法,展现出可控内容创作的巨大潜力。

原文摘要 · Abstract (English)

Recent diffusion-based methods have achieved impressive progress in video content manipulation. However, they typically ignore the accompanying audio, leaving the audio disjointed from the edited results. In this paper, we propose InstructAV2AV, the first end-to-end framework for instruction-guided audio-video joint editing. We first develop a scalable data synthesis pipeline and construct InsAVE-80K, the first large-scale audio-video editing dataset with high-quality source-to-target pairs. With this data foundation, we adapt an audio-video generation backbone to leverage its robust priors. We concatenate the audio-video input with noisy latent codes to anchor the source context, propose the source-instruction gated attention to improve instruction following and content preservation, and introduce a two-stage training strategy to effectively transfer these pre-trained priors. Extensive experiments demonstrate that InstructAV2AV outperforms state-of-the-art methods across 11 metrics spanning three aspects on two evaluation sets, highlighting its potential for controllable content creation. Project page: https://hjzheng.net/projects/InstructAV2AV/.

音视频编辑扩散模型指令控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。