让语音模型会查知识,回答需要背景信息的复杂问题。
Audiopedia: Audio QA with Knowledge
- 构建三类需外部知识的语音问答任务,推动模型理解深度
- 提出音视频实体链接与知识增强框架,显著提升问答准确率
- 适合研究语音理解、多模态推理的学者和开发者
本文提出Audiopedia,一种新型音频问答任务,要求同时具备音频理解与外部知识推理能力。不同于传统仅依赖音频的简单问答,Audiopedia聚焦知识密集型问题,定义三个子任务:(i)单音频问答(s-AQA),基于单一音频样本作答;(ii)多音频问答(m-AQA),需跨多个音频样本进行推理;(iii)检索增强型问答(r-AQA),需检索相关音频以支持回答。我们在大音频语言模型(LALMs)上评估这些任务,发现其表现不佳。为此,提出通用框架,可适配任意LALM,通过音频实体链接(AEL)与知识增强音频大模型(KA2LM)两部分,赋予模型知识推理能力。该工作首次系统探索通过知识密集型任务实现高级音频理解。
原文摘要 · Abstract (English)
In this paper, we introduce Audiopedia, a novel task called Audio Question Answering with Knowledge, which requires both audio comprehension and external knowledge reasoning. Unlike traditional Audio Question Answering (AQA) benchmarks that focus on simple queries answerable from audio alone, Audiopedia targets knowledge-intensive questions. We define three sub-tasks: (i) Single Audio Question Answering (s-AQA), where questions are answered based on a single audio sample, (ii) Multi-Audio Question Answering (m-AQA), which requires reasoning over multiple audio samples, and (iii) Retrieval-Augmented Audio Question Answering (r-AQA), which involves retrieving relevant audio to answer the question. We benchmark large audio language models (LALMs) on these sub-tasks and observe suboptimal performance. To address this, we propose a generic framework that can be adapted to any LALM, equipping them with knowledge reasoning capabilities. Our framework has two components: (i) Audio Entity Linking (AEL) and (ii) Knowledge-Augmented Audio Large Multimodal Model (KA2LM), which together improve performance on knowledge-intensive AQA tasks. To our knowledge, this is the first work to address advanced audio understanding via knowledge-intensive tasks like Audiopedia.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。