首个实现视频到音频高时序对齐的自回归模型,提升音画同步质量。
Temporally Aligned Audio for Video with Autoregression
- 采用高频视觉特征提取与跨模态融合策略,精准捕捉动作细节。
- 在时序对齐和语义相关性上超越现有模型,音频质量相当。
- 适合研究音画生成、多模态对齐的开发者与研究人员。
我们提出 V-AURA,首个在视频到音频生成中实现高时序对齐与语义相关性的自回归模型。V-AURA 使用高帧率视觉特征提取器与跨模态音视频特征融合策略,捕捉细微视觉运动事件并确保精确时序对齐。此外,我们构建了 VisualSound 基准数据集,基于 VGGSound(从 YouTube 提取的真实场景视频)进行筛选,剔除视听不匹配样本。V-AURA 在时序对齐与语义相关性上优于当前最优模型,同时保持相当的音频质量。代码、样例、VisualSound 数据集及模型已公开于 https://v-aura.notion.site。
原文摘要 · Abstract (English)
We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy to capture fine-grained visual motion events and ensure precise temporal alignment. Additionally, we propose VisualSound, a benchmark dataset with high audio-visual relevance. VisualSound is based on VGGSound, a video dataset consisting of in-the-wild samples extracted from YouTube. During the curation, we remove samples where auditory events are not aligned with the visual ones. V-AURA outperforms current state-of-the-art models in temporal alignment and semantic relevance while maintaining comparable audio quality. Code, samples, VisualSound and models are available at https://v-aura.notion.site
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。