测试大模型在多语言干扰下的听觉注意力,发现强单语表现不等于抗干扰能力。
Can Large Audio Language Models Ignore Multilingual Distractors? An Evaluation of Their Selective Auditory Attention Capabilities

- 设计多语言混杂场景评测模型选择性听觉注意力
- 严重信噪比下准确率下降,错误主要来自干扰源混淆
- 分离预处理仍难解决来源误判,易给出自信错误答案
在多语言干扰下实现鲁棒的选择性听觉注意力对大型音频语言模型(LALMs)的可靠部署至关重要。我们提出了MUSA,一个受鸡尾酒会效应启发的多语言基准,用于评估源定位的口语理解与推理能力。每项任务包含一句目标英语对话和一句语义合理的英语、西班牙语、韩语或中文干扰对话,在(1)单语、(2)基于源分离的两阶段、(3)端到端鸡尾酒会三种设置下,于受控信噪比(SNR)条件下评估模型性能。评估了两个闭源和四个开源的LALM,结果发现:良好的单语表现并不保证鲁棒的选择性听觉注意力——在严重低信噪比下,鸡尾酒会准确率显著下降,错误主要源于干扰源引导的来源混淆。此外,分离虽减少声学重叠,但未解决来源归属问题,常导致模型自信地回答错误语音流内容。数据与代码将在发表后公开。
原文摘要 · Abstract (English)
Robust selective auditory attention under multilingual interference is critical for reliable deployment of Large Audio Language Models (LALMs). We introduce MUSA, a cocktail party-inspired multilingual benchmark for source-grounded spoken-language understanding and reasoning. Each item pairs an English target dialogue with a semantically plausible distractor in English, Spanish, Korean, or Chinese, and evaluates models across (1) single, (2) source separation-based two-stage, (3) and end-to-end cocktail party settings under controlled SNRs. Evaluating two closed-source and four open-weight LALMs, we find that strong single performance does not ensure robust selective auditory attention: cocktail party accuracy degrades under severe SNRs, and errors are dominated by distractor-grounded source confusion. In addition, separation reduces acoustic overlap but leaves source attribution unresolved, often yielding confident wrong-stream answers. Data and code will be released upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。