arXiv:2510.08816cs.SDeess.AS2025-10

用非负自编码器分解声音,实现可解释的创意音频操控。

Audible Networks: Deconstructing and Manipulating Sounds with Deep Non-Negative Autoencoders

  • 通过非负约束分解音频,成分可听且对应频谱与时间包络。
  • 多层结构支持从音符到细节的分层分解,精度可调。
  • 适合音乐创作、声音设计者,支持随机化与跨成分合成。

我们提出使用非负自编码器(NAEs)进行声音的可解释性分解与用户引导的声音操控,适用于创意场景。NAEs通过投影梯度下降施加非负性约束,使内部权重和激活值可直接解释为频谱形状与时间包络,且各成分可独立播放为声音事件。多层深度NAE架构支持分层表示,可调节粒度,在多个抽象层次上分解声音:从高阶音符包络到细粒度频谱特征。该框架实现了丰富、可控且可随机化的声学变换,引入了跨成分/跨层合成、分层解构及多种控制音色与事件密度的随机策略。通过可视化与重合成实例,证明了NAEs在基于对象的声音编辑中具备灵活性与可解释性。

原文摘要 · Abstract (English)

We propose the use of Non-Negative Autoencoders (NAEs) for sound deconstruction and user-guided manipulation of sounds for creative purposes. NAEs offer a versatile and scalable extension of traditional Non-Negative Matrix Factorization (NMF)-based approaches for interpretable audio decomposition. By enforcing non-negativity constraints through projected gradient descent, we obtain decompositions where internal weights and activations can be directly interpreted as spectral shapes and temporal envelopes, and where components can themselves be listened to as individual sound events. In particular, multi-layer Deep NAE architectures enable hierarchical representations with an adjustable level of granularity, allowing sounds to be deconstructed at multiple levels of abstraction: from high-level note envelopes down to fine-grained spectral details. This framework enables a wide new range of expressive, controllable, and randomized sound transformations. We introduce novel manipulation operations including cross-component and cross-layer synthesis, hierarchical deconstructions, and several randomization strategies that control timbre and event density. Through visualizations and resynthesis of practical examples, we demonstrate how NAEs can serve as flexible and interpretable tools for object-based sound editing.

音频分解非负矩阵声音编辑深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。