arXiv:2509.04957cs.CVcs.MM2025-09

用多模型映射器实现高效视频转音频,训练量减少84%仍保持优秀效果。

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper

  • 融合双视觉编码器特征,提升语义与时间信息表达
  • 用GPT-2替代线性映射,改善跨模态特征对齐
  • 仅需16%训练规模,性能媲美大规模模型

近期视频转音频(V2A)生成依赖从视频中提取语义和时间特征来引导生成模型。从头训练成本高昂,因此利用基础模型(FMs)进行跨模态知识迁移成为主流。已有工作尝试微调轻量级映射网络,连接预训练视觉编码器与文本到音频生成模型。受此启发,本文提出多基础模型映射器(MFM-Mapper)。相比先前方法,MFM-Mapper通过融合双视觉编码器特征,获得更丰富的语义与时间信息;并以GPT-2取代线性映射,将跨模态特征映射类比为自回归翻译任务,提升特征对齐能力。所提方法在训练效率上表现卓越:仅需先前映射方法16%的训练规模,即可在语义与时间一致性上取得更优结果,且性能媲美在更大规模数据上训练的模型。

原文摘要 · Abstract (English)

Recent Video-to-Audio (V2A) generation relies on extracting semantic and temporal features from video to condition generative models. Training these models from scratch is resource intensive. Consequently, leveraging foundation models (FMs) has gained traction due to their cross-modal knowledge transfer and generalization capabilities. One prior work has explored fine-tuning a lightweight mapper network to connect a pre-trained visual encoder with a text-to-audio generation model for V2A. Inspired by this, we introduce the Multiple Foundation Model Mapper (MFM-Mapper). Compared to the previous mapper approach, MFM-Mapper benefits from richer semantic and temporal information by fusing features from dual visual encoders. Furthermore, by replacing a linear mapper with GPT-2, MFM-Mapper improves feature alignment, drawing parallels between cross-modal features mapping and autoregressive translation tasks. Our MFM-Mapper exhibits remarkable training efficiency. It achieves better performance in semantic and temporal consistency with fewer training consuming, requiring only 16\% of the training scale compared to previous mapper-based work, yet achieves competitive performance with models trained on a much larger scale.

视频转音频多模态高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。