评测大模型对语音音频的多跳推理能力,发现其整合信息困难。
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
- 构建新基准SAKURA,专门评估语音音频的多跳推理。
- 模型虽能提取信息,却难以整合多源内容完成推理。
- 适合关注多模态推理、语音理解的研究者参考。
大型音频语言模型(LALMs)将大语言模型扩展至语音、音频等多模态理解。尽管其在语音与音频处理任务上的表现已广泛研究,但其推理能力仍缺乏系统评估,尤其是多跳推理——即回忆并整合多个事实的能力。现有基准主要关注通用语音处理、对话能力和公平性,忽视了这一关键方面。为填补空白,我们提出SAKURA,一个基于语音与音频信息的多跳推理评测基准。结果表明,即使模型能正确提取相关信息,仍难以将其整合用于多跳推理,揭示了多模态推理中的根本挑战。该研究暴露了LALMs的关键局限,为未来研究提供了洞见与资源。
原文摘要 · Abstract (English)
Large audio-language models (LALMs) extend the large language models with multimodal understanding in speech, audio, etc. While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored. Particularly, their multi-hop reasoning, the ability to recall and integrate multiple facts, lacks systematic evaluation. Existing benchmarks focus on general speech and audio-processing tasks, conversational abilities, and fairness but overlook this aspect. To bridge this gap, we introduce SAKURA, a benchmark assessing LALMs' multi-hop reasoning based on speech and audio information. Results show that LALMs struggle to integrate speech/audio representations for multi-hop reasoning, even when they extract the relevant information correctly, highlighting a fundamental challenge in multimodal reasoning. Our findings expose a critical limitation in LALMs, offering insights and resources for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。