arXiv:2510.00264cs.SDcs.LG2025-10被引 2

2025低资源音频编码挑战赛官方基线系统,兼顾音质与计算效率。

Baseline Systems For The 2025 Low-Resource Audio Codec Challenge

  • 采用卷积神经网络结合残差向量量化,端到端训练
  • 支持在噪声和混响下保持语音透明度与降噪增强能力
  • 适合低算力设备部署,适用于语音编码研究者

低资源音频编码(LRAC)挑战赛旨在推动资源受限环境下神经音频编码技术的发展。首届挑战赛聚焦低资源神经语音编码器,要求在日常噪声和混响条件下稳定运行,并满足严格的计算复杂度、延迟和比特率约束。第1赛道针对透明编码器,目标是在轻度噪声和混响下保持输入语音的感知透明性;第2赛道为增强编码器,融合编码压缩与去噪、去混响功能。本文介绍了2025年LRAC挑战赛两个赛道的官方基线系统。基线模型采用基于残差向量量化(Residual Vector Quantization)的卷积神经编码器,通过对抗损失与重建损失联合训练。文中详述了数据筛选与增强策略、模型架构设计、优化流程及检查点选择标准。

原文摘要 · Abstract (English)

The Low-Resource Audio Codec (LRAC) Challenge aims to advance neural audio coding for deployment in resource-constrained environments. The first edition focuses on low-resource neural speech codecs that must operate reliably under everyday noise and reverberation, while satisfying strict constraints on computational complexity, latency, and bitrate. Track 1 targets transparency codecs, which aim to preserve the perceptual transparency of input speech under mild noise and reverberation. Track 2 addresses enhancement codecs, which combine coding and compression with denoising and dereverberation. This paper presents the official baseline systems for both tracks in the 2025 LRAC Challenge. The baselines are convolutional neural codec models with Residual Vector Quantization, trained end-to-end using a combination of adversarial and reconstruction objectives. We detail the data filtering and augmentation strategies, model architectures, optimization procedures, and checkpoint selection criteria.

音频编码神经编码低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。