发现主流音频模型在频域表示上存在结构性瓶颈,影响音高和音色的可操控性。
Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

- 通过理论分析与实验,揭示卷积编码器在频域上的两个结构限制
- 实际数据中频域原语丢失率达31%-35%,滤波器带宽超出理论极限10-35倍
- 提出轻量级后处理方法,恢复频域局部化表示,提升音高等属性可控性
端到端神经音频模型实现了高质量压缩与生成。尽管表现优异,但其是否真正表征可解释的特征(如音高、音色)尚不明确。这些特征本质上是时间-频率局部化基元的组合。然而,当前主流的步进卷积编码器存在两个可预测的结构性瓶颈:(1) 将基元压缩至混叠等价类,限制表征容量;(2) 降低学习滤波器的频率分辨率,削弱可分性。在真实信号条件下,我们发现基元崩溃率高达31%-35%,滤波器带宽为理论分辨率的10-35倍。为此,我们提出Gabor Latent Refactorization(GLRF),一种无需重训练的轻量级后处理方法,将编码特征重表达为频域局部化基,使滤波器带宽降至理论值的1.5-3倍,同时保持重建保真度,并增强对音高等属性的控制能力。结果表明,现有编码器会系统性地破坏对频域基元的访问,导致相关特征纠缠,而轻量干预可显著恢复可解释性与可控性。
原文摘要 · Abstract (English)
End-to-end neural audio models achieve high-fidelity compression and generation. We might read that performance as evidence they directly represent interpretable features such as pitch and timbre, but a model can produce plausible outputs without doing so. A model may encode these features in any reachable basis, but regardless of which, the features are well described as compositions of time-frequency-localized primitives. Whether state-of-the-art encoders preserve access to these primitives, and thus to compositions of them, remains unclear. Through theoretical analysis and controlled experiments, we show that several state-of-the-art strided convolutional encoders impose two structural bottlenecks, both predictable from architecture and signal structure, on access to these primitives: (1) they collapse primitives into alias equivalence classes, establishing a bound on representational capacity, and (2) they limit the frequency resolution available to learned filters, restricting separability. For well structured data, we find collapse rates of 31-35% and filter bandwidths 10-35x above the theoretical resolution bound, confirming that both bottlenecks arise under realistic signal conditions. We then introduce Gabor Latent Refactorization (GLRF), a lightweight post-hoc intervention that re-expresses encoder latents in a frequency-localized basis, reducing filter bandwidths from 10-35x to 1.5-3x of the theoretical resolution bound while preserving reconstruction fidelity and improving control over attributes like pitch. These results show that the encoders in question predictably degrade access to frequency-localized primitives, entangling the features that depend on them, and that a lightweight, retraining-free intervention can recover much of that access, improving steerability and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。