首个专注音频检索中复杂推理能力的评测基准
ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

- 设计五类推理任务,测试模型对否定、时间顺序等深层语义理解
- 10个主流模型在否定和时长判断上表现差,平均准确率不足50%
- 揭示当前多模态大模型难以通过微调继承推理能力
随着多模态内容快速增长,音频检索已成为媒体搜索、内容组织与智能助手的关键技术。然而,现有评测主要关注语义匹配,未能捕捉真实查询中所需的高级推理能力,如否定理解、时间排序、事件重叠识别和时长区分。为此,我们提出ReasonAudio,首个面向文本-音频检索的推理密集型评测基准,包含1,000个查询和10,000个复合音频片段,覆盖五类基础推理任务:否定(Negation)、顺序(Order)、重叠(Overlap)、时长(Duration)和混合(Mix)。尽管这些任务对人类直观易懂且构建简单,但对当前模型构成重大挑战。对十种先进模型的评估显示:所有模型在推理型音频检索中表现不佳,尤其在否定和时长任务上,准确率普遍低于50%;而在重叠和顺序任务上表现相对较好。此外,基于多模态大语言模型的嵌入模型无法通过对比微调继承其骨干模型的推理能力,表明现有训练范式在检索场景中不足以保留推理性能。
原文摘要 · Abstract (English)
As multimodal content continues to expand at a rapid pace, audio retrieval has emerged as a key enabling technology for media search, content organization, and intelligent assistants. However, most existing benchmarks concentrate on semantic matching and fail to capture the fact that real-world queries often demand advanced reasoning abilities, including negation understanding, temporal ordering, concurrent event recognition, and duration discrimination. To address this gap, we introduce ReasonAudio, the first reasoning-intensive benchmark for Text-Audio Retrieval, comprising 1,000 queries and 10,000 composite audio clips across five fundamental reasoning tasks: Negation, Order, Overlap, Duration, and Mix. Despite their intuitive nature for humans and straightforward construction, these tasks pose significant challenges to current models. Our evaluation of ten state-of-the-art models reveals the following findings: All models struggle with reasoning-intensive audio retrieval, performing particularly poorly on Negation and Duration while showing relatively better results on Overlap and Order. Moreover, Multimodal Large Language Model-based embedding models fail to inherit the reasoning capabilities of their backbones through contrastive fine-tuning, suggesting that current training paradigms are insufficient to preserve reasoning capacity in retrieval settings
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。