arXiv:2602.20981cs.CVcs.AI2026-02中稿 · CVPR被引 2

提出新模型让视频转音频能生成超5分钟长音频,突破长度限制。

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

  • 用分层结构+非因果Mamba网络,支持长时序音频生成。
  • 在长视频转音频任务上实现超5分钟输出,优于现有方法。
  • 训练短视频也能泛化到长音频,无需长数据训练。

视频与音频的多模态对齐面临数据稀缺和文本描述与帧级视频信息不匹配的挑战。本文针对视频到音频生成中的可扩展性问题,研究模型在短样本上训练是否能在测试时泛化到更长序列。为此,提出一种名为MMHNet的增强型多模态分层网络,融合分层结构与非因果Mamba机制,以支持长时序音频生成。实验表明,该方法显著提升长音频生成能力,可生成超过5分钟的音频内容。同时证明,在视频到音频任务中,仅在短实例上训练即可实现长时长测试泛化,无需在长时长数据上进行训练。在长视频转音频基准测试中,所提方法性能优于先前工作,且首次实现超过5分钟的音频生成,而以往方法难以达到此长度。

原文摘要 · Abstract (English)

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling challenge in multimodal-to-audio generation, examining whether models trained on short instances can generalize to longer ones during testing. To tackle this challenge, we present multimodal hierarchical networks so-called MMHNet, an enhanced extension of state-of-the-art video-to-audio models. Our approach integrates a hierarchical method and non-causal Mamba to support long-form audio generation. Our proposed method significantly improves long audio generation up to more than 5 minutes. We also prove that training short and testing long is possible in the video-to-audio generation tasks without training on the longer durations. We show in our experiments that our proposed method could achieve remarkable results on long-video to audio benchmarks, beating prior works in video-to-audio tasks. Moreover, we showcase our model capability in generating more than 5 minutes, while prior video-to-audio methods fall short in generating with long durations.

视频转音频长序列生成Mamba多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。