评测多模态大模型跨模态推理能力,发现顶尖模型仍难应对复杂融合任务。
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
- 构建合成与真实双版本评估集,覆盖文本、图像、音频、视频等多模态组合
- 所有顶级多模态模型在需跨模态整合的任务上表现不佳,平均准确率低于60%
- 适合关注多模态对齐、模型推理缺陷的研究者和开发者
我们提出OmnixR,一个用于评估当前最先进多模态语言模型(如GPT-4o、Gemini)的测评套件。多模态模型需整合文本、视觉、音频等多种信息以完成任务,但现有基准大多仅支持单模态或双模态任务,缺乏对跨模态推理的全面评估。OmnixR包含两个变体:(1) 合成子集,通过自动转换生成包含音频、图像、视频及混合模态的合成数据(Omnify);(2) 真实子集,由专家手工收集与标注的真实场景数据。该评测涵盖视频+音频+文本等多样化组合,提供前所未有的跨模态推理测试环境。实验表明,所有主流多模态模型在需要多模态信息融合的问题上表现欠佳,且分析揭示其推理行为存在显著差异,凸显多模态对齐的挑战。
原文摘要 · Abstract (English)
We introduce OmnixR, an evaluation suite designed to benchmark SoTA Omni-modality Language Models, such as GPT-4o and Gemini. Evaluating OLMs, which integrate multiple modalities such as text, vision, and audio, presents unique challenges. Particularly, the user message might often consist of multiple modalities, such that OLMs have to establish holistic understanding and reasoning across modalities to accomplish the task. Existing benchmarks are limited to single modality or dual-modality tasks, overlooking comprehensive multi-modal assessments of model reasoning. To address this, OmnixR offers two evaluation variants: (1)synthetic subset: a synthetic dataset generated automatically by translating text into multiple modalities--audio, images, video, and hybrids (Omnify). (2)realistic subset: a real-world dataset, manually curated and annotated by experts, for evaluating cross-modal reasoning in natural settings. OmnixR presents a unique evaluation towards assessing OLMs over a diverse mix of modalities, such as a question that involves video, audio, and text, providing a rigorous cross-modal reasoning testbed unlike any existing benchmarks. Our experiments find that all state-of-the-art OLMs struggle with OmnixR questions that require integrating information from multiple modalities to answer. Further analysis highlights differences in reasoning behavior, underscoring the challenges of omni-modal AI alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。