让视频生成长时音频更流畅,修复了拼接和时间错位问题。
LD-LAudio-V1: Video-to-Long-Form-Audio Generation Extension with Dual Lightweight Adapters
- 用双轻量适配器扩展模型,支持长时音频生成。
- 在多个指标上提升超20%,音频更连贯、能量变化更自然。
- 发布纯净标注数据集,适合做音视频生成研究者使用。
从视频生成高质量且时间同步的音频对视频编辑与后期制作至关重要,可为无声视频生成语义一致的音频。现有方法多局限于10秒以内的短片段生成,或依赖含噪数据进行长时生成。为此,我们提出LD-LAudio-V1,作为先进视频到音频模型的扩展,引入双轻量适配器实现长时音频生成。同时发布一个干净、人工标注的视频到音频数据集,仅包含纯净音效,无噪声或伪影。该方法显著减少拼接伪影与时间不一致,同时保持计算高效。相比直接微调短视频,LD-LAudio-V1在多项指标上取得显著提升:$FD_{\text{passt}}$ 450.00 → 327.29(+27.27%),$FD_{\text{panns}}$ 34.88 → 22.68(+34.98%),$FD_{\text{vgg}}$ 3.75 → 1.28(+65.87%),$KL_{\text{panns}}$ 2.49 → 2.07(+16.87%),$KL_{\text{passt}}$ 1.78 → 1.53(+14.04%),$IS_{\text{panns}}$ 4.17 → 4.30(+3.12%),$IB_{\text{score}}$ 0.25 → 0.28(+12.00%),$Energy\Delta10\text{ms}$ 0.3013 → 0.1349(+55.23%),$Energy\Delta10\text{ms(vs.GT)}$ 0.0531 → 0.0288(+45.76%),$Sem.\,Rel.$ 2.73 → 3.28(+20.15%)。该数据集旨在推动长时视频到音频生成研究,已开源至https://github.com/deepreasonings/long-form-video2audio。
原文摘要 · Abstract (English)
Generating high-quality and temporally synchronized audio from video content is essential for video editing and post-production tasks, enabling the creation of semantically aligned audio for silent videos. However, most existing approaches focus on short-form audio generation for video segments under 10 seconds or rely on noisy datasets for long-form video-to-audio zsynthesis. To address these limitations, we introduce LD-LAudio-V1, an extension of state-of-the-art video-to-audio models and it incorporates dual lightweight adapters to enable long-form audio generation. In addition, we release a clean and human-annotated video-to-audio dataset that contains pure sound effects without noise or artifacts. Our method significantly reduces splicing artifacts and temporal inconsistencies while maintaining computational efficiency. Compared to direct fine-tuning with short training videos, LD-LAudio-V1 achieves significant improvements across multiple metrics: $FD_{\text{passt}}$ 450.00 $\rightarrow$ 327.29 (+27.27%), $FD_{\text{panns}}$ 34.88 $\rightarrow$ 22.68 (+34.98%), $FD_{\text{vgg}}$ 3.75 $\rightarrow$ 1.28 (+65.87%), $KL_{\text{panns}}$ 2.49 $\rightarrow$ 2.07 (+16.87%), $KL_{\text{passt}}$ 1.78 $\rightarrow$ 1.53 (+14.04%), $IS_{\text{panns}}$ 4.17 $\rightarrow$ 4.30 (+3.12%), $IB_{\text{score}}$ 0.25 $\rightarrow$ 0.28 (+12.00%), $Energy\Delta10\text{ms}$ 0.3013 $\rightarrow$ 0.1349 (+55.23%), $Energy\Delta10\text{ms(vs.GT)}$ 0.0531 $\rightarrow$ 0.0288 (+45.76%), and $Sem.\,Rel.$ 2.73 $\rightarrow$ 3.28 (+20.15%). Our dataset aims to facilitate further research in long-form video-to-audio generation and is available at https://github.com/deepreasonings/long-form-video2audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。