让语言模型在不听声音的情况下,也能理解声音常识并推理。
AuditoryBench++: Can Language Models Understand Auditory Knowledge without Hearing?
- 用特殊标记和知识注入模拟声音想象,实现文本中推理声音属性。
- 在多项任务上显著超越基础模型和加入声学知识的模型。
- 适合研究多模态推理、声音认知与语言模型能力边界的人看。
人类无需直接听到声音,便可轻松推理音高、响度或声源关联等听觉属性,依赖的是听觉常识。相比之下,语言模型常缺乏此能力,限制了其在多模态交互中的表现。为此,我们提出 AuditoryBench++,一个全面评估文本环境下听觉知识与推理能力的基准。该基准涵盖从基础听觉比较到情境化推理的任务,支持对模型处理和整合听觉概念的细粒度分析。同时,我们引入 AIR-CoT,一种新的听觉想象推理方法,通过特殊标记的跨度检测与知识注入,在推理过程中生成并融合听觉信息。大量实验表明,AIR-CoT 在近期大语言模型与多模态大模型上普遍优于原生模型及加入听觉知识的模型。项目主页见 https://auditorybenchpp.github.io。
原文摘要 · Abstract (English)
Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often lack this capability, limiting their effectiveness in multimodal interactions. As an initial step to address this gap, we present AuditoryBench++, a comprehensive benchmark for evaluating auditory knowledge and reasoning in text-only settings. The benchmark encompasses tasks that range from basic auditory comparisons to contextually grounded reasoning, enabling fine-grained analysis of how models process and integrate auditory concepts. In addition, we introduce AIR-CoT, a novel auditory imagination reasoning method that generates and integrates auditory information during inference through span detection with special tokens and knowledge injection. Extensive experiments with recent LLMs and Multimodal LLMs demonstrate that AIR-CoT generally outperforms both the off-the-shelf models and those augmented with auditory knowledge. The project page is available at https://auditorybenchpp.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。