arXiv:2601.21402eess.AScs.SD2026-01被引 2

让音频生成与编辑在语义空间中进行,提升文本与音频的匹配度。

SemanticAudio: Audio Generation and Editing in Semantic Space

  • 在高层语义空间生成音频,分离语义布局与声学细节。
  • 生成音频与文本描述对齐度显著优于现有方法。
  • 无需重新训练即可实现文本引导的精准属性修改。

近年来,文本到音频生成取得了显著进展,为声音创作者提供了将文字灵感转化为生动音频的强大工具。然而,现有模型主要在变分自编码器(VAE)的声学潜在空间中直接操作,常导致生成音频与文本描述的对齐效果不佳。本文提出SemanticAudio,一种在高层语义空间中进行音频生成与编辑的新框架。该语义空间以紧凑表示捕捉声音事件的全局身份与时间序列,区别于细粒度的声学细节。SemanticAudio采用两阶段流匹配架构:语义规划器首先生成紧凑的语义特征以勾勒全局语义布局,声学合成器随后基于该语义计划生成高保真声学潜在表示。借助这一解耦设计,我们进一步提出一种无需训练的文本引导编辑机制,可在不重新训练的情况下对通用音频实现精确的属性级修改。具体通过源与目标文本提示所对应的速度场差异来引导语义生成轨迹。大量实验表明,SemanticAudio在语义对齐方面超越了现有主流方法。演示地址:https://semanticaudio1.github.io/

原文摘要 · Abstract (English)

In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, existing models predominantly operate directly in the acoustic latent space of a Variational Autoencoder (VAE), often leading to suboptimal alignment between generated audio and textual descriptions. In this paper, we introduce SemanticAudio, a novel framework that conducts both audio generation and editing directly in a high-level semantic space. We define this semantic space as a compact representation capturing the global identity and temporal sequence of sound events, distinct from fine-grained acoustic details. SemanticAudio employs a two-stage Flow Matching architecture: the Semantic Planner first generates these compact semantic features to sketch the global semantic layout, and the Acoustic Synthesizer subsequently produces high-fidelity acoustic latents conditioned on this semantic plan. Leveraging this decoupled design, we further introduce a training-free text-guided editing mechanism that enables precise attribute-level modifications on general audio without retraining. Specifically, this is achieved by steering the semantic generation trajectory via the difference of velocity fields derived from source and target text prompts. Extensive experiments demonstrate that SemanticAudio surpasses existing mainstream approaches in semantic alignment. Demo available at: https://semanticaudio1.github.io/

音频生成语义空间文本编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。