构建多模态多跳推理基准,揭示模型在语音上的显著短板
OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning
- 设计6144个三模态平衡的多跳问题,避免模态捷径
- 开源模型与闭源模型差距大,语音处理性能最弱
- 适合评估多模态推理能力,尤其关注语音模态
多模态大语言模型(MLLMs)已支持文本、视觉和语音的跨模态处理。然而,现有评估框架存在模态捷径和推理路径偏倚等关键缺陷。为此,我们提出OMHBench,一个全新的基准,用于严格评估多模态多跳推理能力。该基准包含6,144个问题,其推理路径在文本、视觉和语音三模态间保持平衡且共同锚定。对13个前沿模型的广泛评估显示:(1) 闭源与开源模型之间存在显著性能差距;(2) 即使闭源模型也对推理路径变化高度敏感,导致多模态接地不对称。尤为突出的是,模型在语音模态上表现明显不佳,凸显了对均衡、多跳多模态智能评估的迫切需求。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evaluation frameworks for such models suffer from critical limitations, including modality shortcuts and biased reasoning paths. To address these challenges, we propose OMHBench, a novel benchmark designed to rigorously evaluate omni-modal multi-hop reasoning. It consists of 6,144 questions with balanced reasoning paths that are jointly grounded across all three modalities. Extensive evaluation of 13 state-of-the-art models reveals that (1) a large performance gap exists between proprietary and open-source MLLMs and (2) even proprietary models exhibit high sensitivity to reasoning path variations, resulting in asymmetric omni-modal grounding. Notably, models struggle when processing the speech modality, underscoring the need for balanced, multi-hop evaluation of omni-modal intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。