测试多模态大模型是否真能做到跨模态一致推理,发现主流模型仍有明显短板。
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
- 设计六种模态组合的测试题,精细诊断模型跨模态表现
- 最强模型Gemini 2.5 Pro在时空推理上准确率不足60%
- 模型对音频输入表现差,且视觉作上下文时一致性更低
全模态大语言模型(OLLMs)旨在统一音频、视觉与文本理解。现有评测主要关注跨模态问答能力,但尚不清楚其是否具备模态无关推理或存在模态偏好。本文提出XModBench,一个大规模三模态基准,专用于衡量跨模态一致性。该基准包含60,828道选择题,涵盖五类任务及所有六种模态组合,可精细诊断模型的模态无关推理、模态差异与方向性不平衡。实验显示,即使最强模型Gemini 2.5 Pro仍面临挑战:(i) 空间与时间推理准确率低于60%;(ii) 当相同语义内容以音频形式呈现时性能显著下降;(iii) 视觉作为上下文时一致性低于文本作为上下文。结果表明当前OLLMs距离真正的模态无关推理仍有很大差距,而XModBench可作为评估和提升跨模态能力的基础工具。所有数据与评测工具将公开于https://xingruiwang.github.io/projects/XModBench/。
原文摘要 · Abstract (English)
Omni-modal large language models (OLLMs) aim to unify audio, vision, and text understanding within a single framework. While existing benchmarks primarily evaluate general cross-modal question-answering ability, it remains unclear whether OLLMs achieve modality-invariant reasoning or exhibit modality-specific biases. We introduce XModBench, a large-scale tri-modal benchmark explicitly designed to measure cross-modal consistency. XModBench comprises 60,828 multiple-choice questions spanning five task families and systematically covers all six modality compositions in question-answer pairs, enabling fine-grained diagnosis of an OLLM's modality-invariant reasoning, modality disparity, and directional imbalance. Experiments show that even the strongest model, Gemini 2.5 Pro, (i) struggles with spatial and temporal reasoning, achieving less than 60% accuracy, (ii) reveals persistent modality disparities, with performance dropping substantially when the same semantic content is conveyed through audio rather than text, and (iii) shows systematic directional imbalance, exhibiting lower consistency when vision serves as context compared to text. These findings indicate that current OLLMs remain far from truly modality-invariant reasoning and position XModBench as a fundamental diagnostic tool for evaluating and improving cross-modal competence. All data and evaluation tools will be available at https://xingruiwang.github.io/projects/XModBench/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。