无需训练即可根据文字指令同步修改音视频内容
Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
- 通过跨模态增量去噪实现音视频协同编辑
- 在AvED-Bench和OAVE数据集上表现优于现有方法
- 适合需要快速音视频内容生成的创作者
本文提出零样本音视频编辑任务,即在不进行额外模型训练的情况下,将原始音视频内容转换为与指定文本提示对齐的内容。为此,我们构建了专为该任务设计的基准数据集AvED-Bench,包含110个10秒长的视频,涵盖来自VGGSound的11个类别,提供多样化的提示与场景,要求音视频元素精准匹配,支持可靠评估。我们发现现有零样本音视频编辑方法在模态间同步性与连贯性方面存在不足,常导致结果不一致。为此,我们提出AvED框架,采用跨模态增量去噪机制,利用音视频交互实现同步且连贯的编辑。在AvED-Bench和最新OAVE数据集上的实验验证了其出色的泛化能力,结果详见https://genjib.github.io/project_page/AVED/index.html。
原文摘要 · Abstract (English)
In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AvED-Bench, designed explicitly for zero-shot audio-video editing. AvED-Bench includes 110 videos, each with a 10-second duration, spanning 11 categories from VGGSound. It offers diverse prompts and scenarios that require precise alignment between auditory and visual elements, enabling robust evaluation. We identify limitations in existing zero-shot audio and video editing methods, particularly in synchronization and coherence between modalities, which often result in inconsistent outcomes. To address these challenges, we propose AvED, a zero-shot cross-modal delta denoising framework that leverages audio-video interactions to achieve synchronized and coherent edits. AvED demonstrates superior results on both AvED-Bench and the recent OAVE dataset to validate its generalization capabilities. Results are available at https://genjib.github.io/project_page/AVED/index.html
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。