arXiv:2505.16774cs.CL2025-05被引 3

首个评估音频大模型指令遵循能力的基准数据集。

IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models

  • 构建包含280个音频-指令-答案对的评测集,覆盖6类结构化要求。
  • 发现现有音频大模型在格式、列表、长度等结构化指令上表现较差。
  • 适合研究多模态大模型、语音理解与指令跟随的学者使用。

大型语言模型(LLM)在文本任务中展现出强大的指令遵循能力,但在与图像或音频等非文本模态对齐后,这种能力常显著下降。尽管已有研究关注文本和视觉-语言模型的指令遵循性能,但基于音频的大语言模型仍缺乏系统评估。为此,我们提出 IFEval-Audio,一个全新的评估数据集,用于衡量音频大模型的指令遵循能力。该数据集包含280个音频-指令-答案三元组,覆盖内容、大小写、符号、列表结构、长度和格式六个维度。每个样本将音频输入与文本指令配对,要求模型生成符合指定结构的输出。我们在多个前沿音频大模型上进行了基准测试,并公开发布该数据集,以推动该新兴领域的研究。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated strong instruction-following capabilities in text-based tasks. However, this ability often deteriorates in multimodal models after alignment with non-text modalities such as images or audio. While several recent efforts have investigated instruction-following performance in text and vision-language models, instruction-following in audio-based large language models remains largely unexplored. To bridge this gap, we introduce IFEval-Audio, a novel evaluation dataset designed to assess the ability to follow instructions in an audio LLM. IFEval-Audio contains 280 audio-instruction-answer triples across six diverse dimensions: Content, Capitalization, Symbol, List Structure, Length, and Format. Each example pairs an audio input with a text instruction, requiring the model to generate an output that follows a specified structure. We benchmark state-of-the-art audio LLMs on their ability to follow audio-involved instructions. The dataset is released publicly to support future research in this emerging area.

音频理解指令遵循多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。