arXiv:2510.23530cs.SDcs.AI2025-10被引 2

让音频自编码器的隐空间变线性,实现直观的音轨混合与调节。

Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization

  • 用数据增强诱导隐空间线性,不改架构或损失函数。
  • 隐空间同时具备可加性和增益等变性,重建质量不变。
  • 适合音乐生成、分离等需要灵活操作音频的场景。

音频自编码器能学习有用的压缩表示,但其非线性隐空间难以进行直观的代数操作(如混合或缩放)。本文提出一种简单训练方法,通过数据增强在高压缩率的一致性自编码器(CAE)中诱导线性特性,从而实现同质性(对标量增益的等变性)和可加性(解码器保持加法性质),且无需改变模型架构或损失函数。训练后,该CAE在编码器和解码器中均表现出线性行为,同时保持重建保真度。我们在音乐源成分和分离任务中测试了该隐空间的实际应用,仅通过简单的隐空间算术即可实现有效操作。本工作提供了一种构建结构化隐空间的直接方法,使音频处理更直观高效。

原文摘要 · Abstract (English)

Audio autoencoders learn useful, compressed audio representations, but their non-linear latent spaces prevent intuitive algebraic manipulation such as mixing or scaling. We introduce a simple training methodology to induce linearity in a high-compression Consistency Autoencoder (CAE) by using data augmentation, thereby inducing homogeneity (equivariance to scalar gain) and additivity (the decoder preserves addition) without altering the model's architecture or loss function. When trained with our method, the CAE exhibits linear behavior in both the encoder and decoder while preserving reconstruction fidelity. We test the practical utility of our learned space on music source composition and separation via simple latent arithmetic. This work presents a straightforward technique for constructing structured latent spaces, enabling more intuitive and efficient audio processing.

音频生成自编码器隐空间线性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。