arXiv:2412.09789cs.SDeess.AS2024-12被引 6

让文本生成音频时可精准控制音量、音高、混响等参数。

SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation

  • 通过解耦音频语义与声学特征,实现参数化控制。
  • 支持音量、音高、混响等7个关键声学参数调节。
  • 不依赖特定模型,适合创意音频制作与设计场景。

文本到音频生成领域虽进展显著,但对生成音频声学特性的精细控制仍研究不足。本文提出SILA方法,通过学习音频语义与声学特征的解耦表示,实现对音量、音高、混响、淡入淡出、亮度、噪声和持续时间等关键声学参数的可控生成,拓展了传统数字信号处理(DSP)的边界。该方法具有模型无关性,能生成高质量且符合用户指定要求的音频,主观与客观评估均验证其有效性,为声音设计与内容创作提供了更丰富、细腻的控制能力。

原文摘要 · Abstract (English)

The field of text-to-audio generation has seen significant advancements, and yet the ability to finely control the acoustic characteristics of generated audio remains under-explored. In this paper, we introduce a novel yet simple approach to generate sound effects with control over key acoustic parameters such as loudness, pitch, reverb, fade, brightness, noise and duration, enabling creative applications in sound design and content creation. These parameters extend beyond traditional Digital Signal Processing (DSP) techniques, incorporating learned representations that capture the subtleties of how sound characteristics can be shaped in context, enabling a richer and more nuanced control over the generated audio. Our approach is model-agnostic and is based on learning the disentanglement between audio semantics and its acoustic features. Our approach not only enhances the versatility and expressiveness of text-to-audio generation but also opens new avenues for creative audio production and sound design. Our objective and subjective evaluation results demonstrate the effectiveness of our approach in producing high-quality, customizable audio outputs that align closely with user specifications.

文本生成音频声学控制音效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。