arXiv:2505.16369cs.SDeess.AS2025-05中稿 · Interspeech 2025被引 13

X-ARES构建音频编码器多任务评测框架,揭示不同领域表现差异。

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance

  • 构建22个任务的综合评测体系,覆盖语音、环境音、音乐等场景
  • 采用线性微调与无参数评估双路径,全面衡量编码器性能
  • 发现主流模型在不同任务中表现差异大,凸显通用音频表征难度

本文提出X-ARES(eXtensive Audio Representation and Evaluation Suite),一个开源基准评测框架,用于系统评估跨多种领域的音频编码器性能。该框架涵盖从语音识别、情感检测到声音事件分类和音乐流派识别等22项任务,涉及语音、环境音和音乐三大领域。通过线性微调和无参数评估两种方式,对先进音频编码器进行广泛评测,结果揭示了模型在不同任务与领域间存在显著性能差异,凸显通用音频表示学习的复杂性。

原文摘要 · Abstract (English)

We introduces X-ARES (eXtensive Audio Representation and Evaluation Suite), a novel open-source benchmark designed to systematically assess audio encoder performance across diverse domains. By encompassing tasks spanning speech, environmental sounds, and music, X-ARES provides two evaluation approaches for evaluating audio representations: linear fine-tuning and unparameterized evaluation. The framework includes 22 distinct tasks that cover essential aspects of audio processing, from speech recognition and emotion detection to sound event classification and music genre identification. Our extensive evaluation of state-of-the-art audio encoders reveals significant performance variations across different tasks and domains, highlighting the complexity of general audio representation learning.

音频编码评测框架多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。