评测语音语言模型的指令遵循能力与灾难性遗忘问题
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
- 构建Speech-IFeval框架,分离评估语音模型的指令遵循与语音感知能力
- 多数语音语言模型在基础指令上表现远差于纯文本大模型,且对提示敏感
- 适合关注多模态模型鲁棒性与评估方法的研究者
我们提出Speech-IFeval评估框架,用于衡量语音感知语言模型(SLMs)的指令遵循能力并量化其灾难性遗忘程度。近期的SLMs将语音感知与大语言模型(LLMs)融合,常因以语音为中心的训练导致文本能力下降。现有基准混淆了语音感知与指令遵循任务,阻碍了对这两种能力的独立评估。为此,我们提供了一个诊断SLM指令遵循能力的基准。研究发现,大多数SLMs在基础指令上表现不佳,显著落后于纯文本大模型;同时,这些模型对提示变化极为敏感,输出不一致且不可靠。我们揭示了核心挑战,并为未来研究提供指导,强调评估应超越任务级指标。
原文摘要 · Abstract (English)
We introduce Speech-IFeval, an evaluation framework designed to assess instruction-following capabilities and quantify catastrophic forgetting in speech-aware language models (SLMs). Recent SLMs integrate speech perception with large language models (LLMs), often degrading textual capabilities due to speech-centric training. Existing benchmarks conflate speech perception with instruction-following, hindering evaluation of these distinct skills. To address this gap, we provide a benchmark for diagnosing the instruction-following abilities of SLMs. Our findings show that most SLMs struggle with even basic instructions, performing far worse than text-based LLMs. Additionally, these models are highly sensitive to prompt variations, often yielding inconsistent and unreliable outputs. We highlight core challenges and provide insights to guide future research, emphasizing the need for evaluation beyond task-level metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。