构建开放音频编码评测基准,解决评估不统一问题
OpenACE: An Open Benchmark for Evaluating Audio Coding Performance
- 提出包含多样化内容的全频段音频编码评测基准
- 对比了Opus、EVS、LC3等编码器在不同场景下的质量表现
- 适合音频编码研究者和开发者用于公平性能对比
音频与语音编码缺乏统一的评估标准和开源测试工具。许多候选系统在专有、不可复现或小规模数据集上评估,而基于机器学习的编码器常在与训练数据分布相似的数据上测试,导致对传统数字信号处理编码器不公平。本文提出一个包含多样内容类型的全频段音频与语音编码质量评测基准,包含传统开源测试向量。以Opus、3GPP EVS及ETSI LC3/LC3+(用于蓝牙低功耗音频)为例,展示编码质量评估的应用。此外,还分析了16 kbps下情感语音编码的质量差异。该开源基准有助于推动音频编码的民主化,代码已公开于https://github.com/JozefColdenhoff/OpenACE。
原文摘要 · Abstract (English)
Audio and speech coding lack unified evaluation and open-source testing. Many candidate systems were evaluated on proprietary, non-reproducible, or small data, and machine learning-based codecs are often tested on datasets with similar distributions as trained on, which is unfairly compared to digital signal processing-based codecs that usually work well with unseen data. This paper presents a full-band audio and speech coding quality benchmark with more variable content types, including traditional open test vectors. An example use case of audio coding quality assessment is presented with open-source Opus, 3GPP's EVS, and recent ETSI's LC3 with LC3+ used in Bluetooth LE Audio profiles. Besides, quality variations of emotional speech encoding at 16 kbps are shown. The proposed open-source benchmark contributes to audio and speech coding democratization and is available at https://github.com/JozefColdenhoff/OpenACE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。