arXiv:2510.23558cs.SDcs.CL2025-10被引 6

评测大模型对指令变化的敏感度,发现顶尖模型也易出错。

ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models

  • 构建三维度动态基准,测试指令描述、输出格式和任务组合的影响
  • 顶尖音频大模型在指令微调后性能下降超30%,体现严重敏感性
  • 适合研究指令鲁棒性或落地应用的AI开发者参考

大型音频语言模型(LALMs)结合声学感知与大语言模型(LLMs),从音频中提取并理解多样化信息,受到学术界和产业界的广泛关注。然而,现有LALMs对指令表述高度敏感,影响指令遵循率和任务表现。现有基准缺乏系统性评估。本文提出ISA-Bench,一个动态基准,从指令描述、输出格式和任务构成三个维度评估LALMs的指令敏感性。我们使用该基准评估近期开源与专有LALMs,分析其在受控指令变化下的合规性与准确性。实验结果表明,即使最先进的LALMs也存在显著指令敏感性,导致基础音频理解任务性能下降。为缓解此问题,我们在特定构造的复杂指令变体数据集上微调Qwen2-Audio,显著提升指令遵循能力。但这也引发非平凡的灾难性遗忘:模型在适应新指令风格时丧失部分原有任务能力。该基准为评估和改进LALMs的指令鲁棒性提供标准化基础,强调真实场景中构建指令鲁棒音频理解系统的必要性。

原文摘要 · Abstract (English)

Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted intense interest from both academic and industrial communities. However, existing LALMs are highly sensitive to how instructions are phrased, affecting both (i) instruction-following rates and (ii) task performance. Yet, no existing benchmarks offer a systematic and comprehensive evaluation of this sensitivity. We introduce ISA-Bench, a dynamic benchmark evaluating instruction sensitivity for LALMs along three axes: instruction description, output format, and task composition. We assess recent open-source and proprietary LALMs using ISA-Bench, profiling both compliance and accuracy under controlled instruction variations. Experimental results reveal that even state-of-the-art LALMs suffer significant instruction sensitivity, leading to degraded performance on fundamental audio understanding tasks. To mitigate this issue, we fine-tune Qwen2-Audio on a specifically constructed complex instruction-variant dataset, achieving a marked improvement in instruction-following performance. However, this also induces nontrivial catastrophic forgetting: the model loses some previously mastered task capabilities when exposed to new instruction styles. Our benchmark provides a standardized basis for assessing and improving instruction sensitivity in LALMs, underscoring the need for instruction-robust audio understanding in real-world pipelines.

音频大模型指令敏感性Qwen2-Audio基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。