用语言-音频蒸馏实现低帧率高保真语义音频压缩
SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation
- 在频域构建轻量级变分自编码器,结合对比学习与CLAP嵌入蒸馏
- 7.8赫兹极低潜码率下仍保持顶级音质,分类任务表现优于同类模型
- 可零样本生成音频描述,适合多模态应用与资源受限场景
现代生成与多模态模型越来越依赖紧凑的潜在表示,在语义丰富性与高保真重建之间取得平衡。我们提出SALAD-VAE,一种连续且高度紧凑的语义音频变分自编码器,工作于频域,以极低的潜码率(7.8 Hz)实现最先进的压缩效果,同时揭示语义结构并生成高质量音频。通过增强标准VAE的语义损失与数据增强,特别是对比学习和基于CLAP的嵌入蒸馏,使其能跨多样音频领域泛化。相比同类先进VAE,SALAD-VAE架构显著更轻量,计算复杂度更低,重建质量相当,但在多种分类基准上持续领先。此外,所提出的额外损失函数训练出一个可零样本用于音频描述与分类的CLAP投影层,匹配预训练的CLAP音文嵌入。
原文摘要 · Abstract (English)
Modern generative and multimodal models increasingly rely on compact latent representations that trade and balance semantic richness with high-fidelity reconstruction. We introduce SALAD-VAE, a continuous and highly compact semantic Audio Variational Autoencoder, which operates in the frequency domain and achieves state-of-the-art compression with very low latent frame rate (7.8 Hz) while surfacing semantic structure and producing high audio quality. We enhance the standard VAE semantic losses and augmentation, specifically contrastive learning and CLAP-based embedding distillation, enabling it to generalize across diverse audio domains. With a significantly less computational complex architecture than comparable state-of-the-art VAEs, SALAD-VAE matches their reconstruction quality while it consistently outperforms them on a wide range of classification benchmarks. Furthermore, the proposed additional loss function provides a trained CLAP projection layer, which can be used zero-shot audio captioning and classification matching pretrained CLAP audio-text embeddings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。