arXiv:2603.22728cs.SDeess.AS2026-03被引 3

为大模型音频编码器设立标准化评测基准,推动多模态语言模型发展

The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models

  • 构建统一生成评估框架XARES-LLM,测试编码器在多种任务中的表现
  • 通过解耦编码器与大模型微调,实现通用音频表征的标准化评测
  • 适合研究音频表征、多模态大模型的开发者和评测人员参考

本文提出Interspeech 2026音频编码器能力挑战赛,旨在评估并推进预训练音频编码器作为大型音频语言模型(LALMs)前端模块的性能。尽管LALMs已展现出对复杂声学场景的出色理解能力,但其表现依赖于底层音频编码器所提取表征的语义丰富度。该挑战通过提供统一的生成式评估框架XARES-LLM,从多样化的下游分类与生成任务中评估提交的编码器。通过将编码器开发与大语言模型微调解耦,该挑战建立了通用音频表征的标准协议,可有效服务于下一代多模态语言模型。

原文摘要 · Abstract (English)

This paper presents the Interspeech 2026 Audio Encoder Capability Challenge, a benchmark specifically designed to evaluate and advance the performance of pre-trained audio encoders as front-end modules for Large Audio Language Models (LALMs). While LALMs have shown remarkable understanding of complex acoustic scenes, their performance depends on the semantic richness of the underlying audio encoder representations. This challenge addresses the integration gap by providing a unified generative evaluation framework, XARES-LLM, which assesses submitted encoders across a diverse suite of downstream classification and generation tasks. By decoupling encoder development from LLM fine-tuning, the challenge establishes a standardized protocol for general-purpose audio representations that can effectively be used for the next generation of multimodal language models.

音频编码大模型评测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。