arXiv:2606.07182eess.AS2026-06

让视频生成音频时可独立控制音色和节奏,效果更自然。

Audio Imitator: Controlling Timbre and Tempo in Video2Audio Synthesis with Audio Reference

论文配图:Audio Imitator: Controlling Timbre and Tempo in Video2Audio Synthesis with Audio Reference
图 1 · 摘自论文原文
  • 将音色和节奏拆解为独立控制变量,取代整体参考音频
  • 在VGGSound上音色相似度提升18.3%,且保持语义对齐
  • 适合需要精细音频风格调节的创作与影视制作

视频转音频生成已实现从无声视频中获得语义一致性和时间同步的音频。然而,音频包含丰富的风格属性如音色和节奏,仅靠视觉或文本输入难以推断。虽然参考音频可作为额外条件,但通常被当作整体信号,限制了细粒度风格控制。我们提出AudioIM,一种属性感知框架,将音色和节奏显式建模为独立控制因素,而非依赖整体提示条件。双编码器分别提取互补的音色相关与节奏相关表征,并通过全局条件注入。基于掩码的训练策略实现了推理阶段有效的潜在提示条件化。在VGGSound数据集上的实验表明,该方法在保持语义对齐与同步的前提下,提升了音色相似度(+18.3%),音频样本可访问:https://anonymousdemo757.github.io/。

原文摘要 · Abstract (English)

Video-to-audio generation has made significant progress in achieving semantic consistency and temporal alignment from silent videos. However, audio contains rich stylistic attributes such as timbre and tempo that are difficult to infer from visual and textual inputs alone. While reference audio can serve as additional conditioning, it is typically treated as a holistic signal, limiting fine-grained style control. We propose AudioIM, an attribute-aware framework that explicitly models timbre and tempo as separate control factors rather than relying on holistic prompt conditioning. Dual encoders extract complementary timbre-related and tempo-related representations, which are injected through global conditioning. A masking-based training strategy enables effective latent prompt conditioning at inference. Experiments on VGGSound show improved style similarity while preserving semantic alignment and synchronization. Audio samples are available at: https://anonymousdemo757.github.io/.

音频生成风格控制视频转音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。