PhoenixCodec在极低资源下实现高质量语音编码,支持1kbps与6kbps双速率。
PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
- 采用非对称频时架构,优化资源分配以降低计算开销。
- 在1kbps下语音可懂度最优,噪声与混响场景表现领先。
- 适合极端低资源环境,如物联网语音传输与边缘设备部署。
本文提出PhoenixCodec,一种专为极低资源条件设计的神经语音编码与解码框架。系统集成优化的非对称频时架构、循环校准与精炼(CCR)训练策略,以及抗噪微调流程。在严格约束下——计算量低于700 MFLOPs,延迟小于30毫秒,支持1 kbps与6 kbps双速率——现有方法面临效率与质量的权衡。PhoenixCodec通过缓解传统解码器的资源分散问题,利用CCR提升优化稳定性,并通过含噪样本微调增强鲁棒性。在LRAC 2025挑战赛第一赛道中,该系统位列第三,且在1 kbps下于真实噪声与混响场景及纯净测试中均取得最佳可懂度表现,验证其有效性。
原文摘要 · Abstract (English)
This paper presents PhoenixCodec, a comprehensive neural speech coding and decoding framework designed for extremely low-resource conditions. The proposed system integrates an optimized asymmetric frequency-time architecture, a Cyclical Calibration and Refinement (CCR) training strategy, and a noise-invariant fine-tuning procedure. Under stringent constraints - computation below 700 MFLOPs, latency less than 30 ms, and dual-rate support at 1 kbps and 6 kbps - existing methods face a trade-off between efficiency and quality. PhoenixCodec addresses these challenges by alleviating the resource scattering of conventional decoders, employing CCR to enhance optimization stability, and enhancing robustness through noisy-sample fine-tuning. In the LRAC 2025 Challenge Track 1, the proposed system ranked third overall and demonstrated the best performance at 1 kbps in both real-world noise and reverberation and intelligibility in clean tests, confirming its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。