arXiv:2503.10700cs.CVcs.MM2025-03被引 7

用文本引导视频转音频,让生成音乐更贴切画面。

TA-V2A: Textually Assisted Video-to-Audio Generation

  • 融合语言模型增强视频语义理解,提升特征表达
  • 通过文本调制扩散模型,实现高效精准音频生成
  • 适合需要精准音画匹配的多媒体创作场景

随着人工智能生成内容(AIGC)的发展,视频转音频(V2A)成为多媒体编辑、增强现实和自动化内容创作中的关键方向。尽管基于Transformer和扩散模型的音频生成技术已取得进展,但现有模型仅依赖帧级特征,难以保留视频的时序语义信息。为此,我们提出TA-V2A,通过整合语言、音频与视频特征,提升潜在空间中的语义表征能力。利用大语言模型增强视频理解,并借助文本引导实现语义丰富表达。基于扩散模型的系统采用自动文本调制机制,提高推理质量和效率,支持个性化文本控制接口。该方法在保持时间对齐的同时增强了语义表达,显著提升了视频到音频生成的准确性和连贯性。

原文摘要 · Abstract (English)

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While Transformer and Diffusion models have advanced audio generation, a significant challenge persists in extracting precise semantic information from videos, as current models often lose sequential context by relying solely on frame-based features. To address this, we present TA-V2A, a method that integrates language, audio, and video features to improve semantic representation in latent space. By incorporating large language models for enhanced video comprehension, our approach leverages text guidance to enrich semantic expression. Our diffusion model-based system utilizes automated text modulation to enhance inference quality and efficiency, providing personalized control through text-guided interfaces. This integration enhances semantic expression while ensuring temporal alignment, leading to more accurate and coherent video-to-audio generation.

视频转音频文本引导扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。