arXiv:2608.08667eess.AS2026-08

提出统一框架,解析音频生成中表示与建模的匹配关系。

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies

论文配图:A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies
图 1 · 摘自论文原文
  • 从表示设计到建模策略,构建双维度评估体系
  • 揭示不同架构在依赖范围与不确定性上的权衡机制
  • 适合研究音频生成系统设计的学者参考

每种音频生成系统都需做出两个相互关联的决策:生成何种表示,以及如何建模其分布。本文围绕这一耦合关系组织音频生成建模。在表示设计方面,通过四个目标(表示负担、失真、模型可实现性、流式兼容性)对比离散、连续和混合潜变量。在分布建模方面,不将潜变量的复杂性视为单一标量,而是引入两个诊断维度:依赖范围(有用上下文延伸距离)和条件模糊性(条件后仍存的不确定性)。该框架细化了语义-声学直觉:长依赖范围变量应具备全局建模能力;条件模糊的细节可交由局部或迭代生成器处理。应用于代表性系统表明:RVQ 的残差顺序提供有序容量而非有序语义;AudioLM 的语义-声学级联仅为边界的一种显式设置,并非普适模板;自回归、迭代优化与混合设计的核心差异在于依赖范围与关键路径生成成本之间的权衡。离散与连续潜变量的区别体现在输出接口;而依赖范围、条件模糊性和流式处理决定该接口应如何建模。本文不罗列具体系统,而是提供比较表示-建模对的评估与设计框架。

原文摘要 · Abstract (English)

Every audio generative system makes two coupled decisions: what representation to generate, and how to model its distribution. This paper organizes audio generative modeling around this coupling. For representation design, we compare discrete, continuous, and hybrid latents through four objectives: representation burden, distortion, empirical modelability, and streaming compatibility. For distribution modeling, rather than treating a latent's difficulty as an intrinsic scalar, we use two diagnostic dimensions: dependency horizon, how far useful context extends, and conditional ambiguity, how much uncertainty remains after conditioning. These dimensions refine the common semantic-versus-acoustic intuition: variables with a long dependency horizon should receive global modeling capacity. Conditionally ambiguous detail may be delegated to a local or iterative generator. Applied to representative systems, this view shows that RVQ's residual order gives ordered capacity but not ordered semantics, that AudioLM's semantic-versus-acoustic cascade is one explicit placement of this boundary rather than a universal template, and that autoregression, iterative refinement, and hybrid designs differ chiefly in how they trade dependency horizon against critical-path generation cost. The distinction between discrete and continuous latents describes the output interface; dependency horizon, conditional ambiguity, and streaming determine how that interface should be modeled. Rather than cataloguing individual systems, we provide an evaluation and design framework for comparing representation-model pairs.

音频生成潜变量建模框架生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。