解决长视频生成配乐不连贯问题,提升音画同步效果。
LoVA: Long-form Video-to-Audio Generation
- 基于扩散Transformer架构,实现长视频到音频的端到端生成。
- 在10秒短视频上表现相当,在长视频任务中显著优于现有方法。
- 适合影视后期、长视频内容创作等需要高质量配乐的场景。
视频到音频(V2A)生成对视频编辑与后期处理至关重要,可为无声视频生成语义对齐的音频。然而,现有方法大多聚焦于生成时长小于10秒的短音频片段,对长视频输入关注不足。当前基于UNet的扩散模型在处理长音频生成时,常出现最终拼接音频中的不一致性问题。本文首次强调长视频V2A的重要性,并提出LoVA——一种面向长视频的新型生成模型。基于扩散Transformer(DiT)架构,LoVA在生成长音频方面优于现有自回归模型和基于UNet的扩散模型。大量客观与主观实验表明,LoVA在10秒基准测试中表现相当,而在长视频输入基准测试中全面超越所有基线模型。
原文摘要 · Abstract (English)
Video-to-audio (V2A) generation is important for video editing and post-processing, enabling the creation of semantics-aligned audio for silent video. However, most existing methods focus on generating short-form audio for short video segment (less than 10 seconds), while giving little attention to the scenario of long-form video inputs. For current UNet-based diffusion V2A models, an inevitable problem when handling long-form audio generation is the inconsistencies within the final concatenated audio. In this paper, we first highlight the importance of long-form V2A problem. Besides, we propose LoVA, a novel model for Long-form Video-to-Audio generation. Based on the Diffusion Transformer (DiT) architecture, LoVA proves to be more effective at generating long-form audio compared to existing autoregressive models and UNet-based diffusion models. Extensive objective and subjective experiments demonstrate that LoVA achieves comparable performance on 10-second V2A benchmark and outperforms all other baselines on a benchmark with long-form video input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。