现有音频语言模型难区分声音存在与不存在,导致误判。
The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

- 构建新评测框架NegEval-Audio,检测模型对声音缺失的感知能力。
- 在两个数据集上,否定类任务准确率远低于随机水平。
- 模型普遍存在肯定偏见,需专门训练提升对否定的理解。
音频-语言嵌入模型如CLAP广泛用于匹配存在的声音事件,却很少评估其对否定的处理能力。我们发现,仅基于肯定的评估掩盖了关键缺陷:这些模型无法区分肯定与否定描述,将两者映射到几乎相同的向量表示。为此,我们提出NegEval-Audio框架,将现有数据集转换为两种否定感知任务——检索否定(Retrieval-Neg)和多选否定(MCQ-Neg),以检验模型是否能识别缺失事件。在AudioCaps和Clotho数据集上,模型在否定任务中表现急剧下降,否定型多选任务准确率远低于随机水平,且这种缺陷在近期基于多模态大模型的嵌入方法中依然存在。尽管无训练调优方法可轻微改善多选任务表现,但在检索任务中提升有限。这表明肯定偏见是表征几何中的根本性缺陷,亟需引入显式的否定感知训练目标。
原文摘要 · Abstract (English)
Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated captions to nearly identical representations. To expose this blind spot, we introduce NegEval-Audio, a framework that converts existing datasets into two negation-aware tasks, Retrieval-Neg and Multiple-Choice Negation (MCQ-Neg), to probe whether models distinguish present from absent events. On AudioCaps and Clotho, performance degrades sharply under negation, with negation-type MCQ accuracy falling far below chance, and the failure persists even for a recent multimodal LLM-based embedding model. While a training-free steering method improves MCQ-Neg, it yields marginal gains for Retrieval-Neg. This indicates that affirmation bias is a fundamental flaw in the representation geometry, necessitating explicit negation-aware training objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。