arXiv:2503.16853cs.CLcs.AI2025-03ACL被引 2

用生成模型动态构建听觉知识,提升语言模型的音频常识理解能力

Imagine to Hear: Auditory Knowledge Generation can be an Effective Assistant for Language Models

  • 通过生成模型从文本中动态生成听觉知识,无需外部音频库
  • 在AuditoryBench上达到当前最佳性能,准确率显著优于基线方法
  • 适合需要音频常识推理的语言模型研究者和开发者

仅基于文本预训练的语言模型在需要听觉常识的任务上表现不佳。以往方法依赖外部音频数据库检索知识,存在相关音频缺失和构建成本高的问题。为此,我们提出Imagine to Hear,一种基于生成模型动态构建听觉知识的新方法。该框架从提示文本中检测多个与音频相关的片段,并生成对应听觉知识。我们设计了基于CLAP的过滤采样器和语言-音频融合模块,以高效处理多源听觉知识。实验表明,该方法在不依赖外部数据库的情况下,在AuditoryBench上达到当前最优性能,验证了生成式方法的有效性。

原文摘要 · Abstract (English)

Language models pretrained on text-only corpora often struggle with tasks that require auditory commonsense knowledge. Previous work addresses this problem by augmenting the language model to retrieve knowledge from external audio databases. This approach has several limitations, such as the potential lack of relevant audio in databases and the high costs associated with constructing the databases. To address these issues, we propose Imagine to Hear, a novel approach that dynamically generates auditory knowledge using generative models. Our framework detects multiple audio-related textual spans from the given prompt and generates corresponding auditory knowledge. We develop several mechanisms to efficiently process multiple auditory knowledge, including a CLAP-based rejection sampler and a language-audio fusion module. Our experiments show that our method achieves state-of-the-art performance on AuditoryBench without relying on external databases, highlighting the effectiveness of our generation-based approach.

语言模型听觉常识生成模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。