arXiv:2602.15749cs.SDeess.AS2026-02被引 4

新音频自编码器提升压缩比与速度,支持多种格式统一建模。

A Generative-First Neural Audio Autoencoder

  • 以生成为导向设计,统一处理连续与离散潜在表示
  • 编码速度提升10倍,潜变量率降低1.6倍,时序下采样达3360倍
  • 单模型支持多通道音频格式,适合低延迟生成应用

神经自编码器是生成模型的核心。实际大规模使用神经自编码器进行生成建模需要快速编码、低潜变量率以及跨表示的统一模型。现有方法以重建为先,导致潜变量率高、编码慢,并且需为离散与连续潜变量、不同音频通道格式分别设计独立架构,阻碍了从预处理到推理条件化的流程。本文提出一种生成导向的音频自编码架构,将时序下采样率从2048倍提升至3360倍,支持连续与离散表示及常见音频通道格式于同一模型中。通过平衡压缩率、质量与速度,该模型实现10倍编码加速、1.6倍率降低,消除通道格式特异性变体,同时保持竞争力的重建质量。例如,60秒单声道信号可压缩为788个标记,使此前受限于计算成本的生成建模更可行。

原文摘要 · Abstract (English)

Neural autoencoders underpin generative models. Practical, large-scale use of neural autoencoders for generative modeling necessitates fast encoding, low latent rates, and a single model across representations. Existing approaches are reconstruction-first: they incur high latent rates, slow encoding, and separate architectures for discrete vs. continuous latents and for different audio channel formats, hindering workflows from preprocessing to inference conditioning. We introduce a generative-first architecture for audio autoencoding that increases temporal downsampling from 2048x to 3360x and supports continuous and discrete representations and common audio channel formats in one model. By balancing compression, quality, and speed, it delivers 10x faster encoding, 1.6x lower rates, and eliminates channel-format-specific variants while maintaining competitive reconstruction quality. This enables applications previously constrained by processing costs: a 60-second mono signal compresses to 788 tokens, making generative modeling more tractable.

音频生成自编码器压缩效率统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。