测试大模型从音频示例中学习任务模式的能力,发现其只懂格式不懂实质。
ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability
- 构建三阶段框架,逐步减少文字提示,评估音频条件下的上下文学习能力
- 六种模型在四类任务中表现:格式合规性提升,但核心任务性能反而下降
- 适合关注多模态模型泛化与理解缺陷的研究者
尽管大型音频-语言模型(LALMs)已被证实指令遵循能力下降,但在音频条件下从上下文示例中推断任务模式的能力仍缺乏研究。为填补这一空白,我们提出ALICE,一个三阶段框架,通过逐步减少文本引导,系统评估LALMs在音频条件下的上下文学习能力。我们在四种音频理解任务上,对六种LALMs进行评估,涵盖两类输出约束。结果揭示出所有阶段和模型的一致不对称现象:上下文示范能显著提升格式合规性,但无法提升甚至损害核心任务性能。这表明LALMs可从示范中提取表面格式模式,却难以利用跨模态语义关联,从音频条件示例中可靠推断任务目标,凸显当前跨模态融合的潜在局限。
原文摘要 · Abstract (English)
While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied. To address this gap, we present ALICE, a three-stage framework that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability under audio conditioning. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across all stages and LALMs: in-context demonstrations reliably improve format compliance but fail to improve, and often degrade, the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from audio-conditioned examples, highlighting potential limitations in current cross-modal integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。