arXiv:2503.01485cs.SDcs.LG2025-03ICLR被引 30

FlowDec用新方法实现低码率高保真音频编码,比之前更自然。

FlowDec: A flow-based full-band general audio codec with high perceptual quality

  • 用新型条件流匹配+非对抗训练,实现4kbps低码率编码
  • 在4kbps下音质超越传统GAN模型DAC,听感接近最佳水平
  • 只需6次后处理计算,大幅降低推理开销,适合实际部署

我们提出FlowDec,一种针对48 kHz采样通用音频的神经全频段编码器,结合非对抗训练与基于新型条件流匹配的随机后滤波器。相比基于得分匹配的ScoreDec,将适用范围从语音扩展至通用音频,码率从24 kbit/s降至最低4 kbit/s,同时提升输出质量,并将所需后滤波器深度神经网络评估次数从60次减少至6次,无需微调或蒸馏。我们提供了理论分析与几何直观,对比ScoreDec及另一近期使用流匹配的工作,并对所提组件进行消融研究。结果显示,FlowDec是当前主流GAN驱动神经编码器的有力替代方案,在FAD评分上优于成熟GAN编码器DAC,听觉测试表现相当,且在语音和音乐的谐波结构重建上更自然。

原文摘要 · Abstract (English)

We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared to the prior work ScoreDec which is based on score matching, we generalize from speech to general audio and move from 24 kbit/s to as low as 4 kbit/s, while improving output quality and reducing the required postfilter DNN evaluations from 60 to 6 without any fine-tuning or distillation techniques. We provide theoretical insights and geometric intuitions for our approach in comparison to ScoreDec as well as another recent work that uses flow matching, and conduct ablation studies on our proposed components. We show that FlowDec is a competitive alternative to the recent GAN-dominated stream of neural codecs, achieving FAD scores better than those of the established GAN-based codec DAC and listening test scores that are on par, and producing qualitatively more natural reconstructions for speech and harmonic structures in music.

音频编码流模型低码率神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。