高压缩比音乐音频编码器,兼顾重建质量与生成性能。
SAME: A Semantically-Aligned Music Autoencoder
- 用Transformer骨干+语义正则化实现4096倍时间压缩
- 保持高质量重建和下游生成效果,计算开销显著降低
- 提供大模型与可部署小模型,支持开源使用
潜在表示是现代生成模型的核心。在音频领域,它们通常由神经音频编解码器自动编码器生成。本文提出SAME(语义对齐音乐自编码器),一种用于立体声音乐和通用音频的自编码器,在保持重建质量和下游生成性能的同时,实现了4096×的时间压缩比。通过结合基于Transformer的主干网络、一组语义正则化方法、相位感知重建损失以及改进的判别器设计,该架构在提升压缩率的同时,依赖优化良好的Transformer原语大幅降低计算成本。本文发布了两个版本(大型SAME-L和可在CPU部署的SAME-S),均以开源权重形式开放。
原文摘要 · Abstract (English)
Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio-codec autoencoder. In this work we introduce SAME (Semantically-Aligned Music autoEncoder), an autoencoder for stereo music and general audio that reaches a 4096$\times$ temporal compression ratio while maintaining reconstruction quality and downstream generative performance. We achieve this by combining a tranformer-based backbone with set of semantic regularisation approaches, phase-aware reconstruction losses and improved discriminator designs. The architecture delivers substantial computational cost benefits, through both its high compression ratio and its reliance on well-optimised transformer primitives. Two variants (a large SAME-L and a CPU-deployable SAME-S) are released in open-weights form.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。