arXiv:2409.15272cs.CLcs.AI2024-09NeurIPS被引 95

构建三模态统一评估基准,揭示现有模型在跨模态推理上的严重不足。

OmniBench: Towards The Future of Universal Omni-Language Models

  • 提出OmniBench基准,评测视觉、听觉与文本的联合理解能力
  • 开源模型在三模态任务中准确率普遍低于50%,指令遵循能力弱
  • 发布84.5万样本指令数据集,推动三模态模型训练发展

近年来多模态大语言模型(MLLMs)致力于整合与解析跨模态数据,但其同时处理与推理多种模态的能力仍缺乏系统评估,主要受限于缺乏全面的模态专用基准。我们提出OmniBench,一个新型基准,用于严格评估模型在视觉、听觉与文本输入上同时进行识别、理解与推理的能力。能实现三模态处理的语言模型称为全语言模型(OLMs)。OmniBench以高质量人工标注为特色,确保正确回答需融合三种模态的综合理解。主要发现显示:(1)开源OLMs在三模态情境下的指令遵循与推理能力存在明显缺陷;(2)多数基线模型即使获得图像或音频的替代文本表示,准确率也普遍低于50%。这表明现有MLLM训练范式常忽略从文本、图像和音频构建一致语境的能力。为此,我们构建了包含84.5万样本的指令调优数据集OmniInstruct,用于训练适应三模态环境的OLMs。我们呼吁未来研究应聚焦更鲁棒的三模态融合技术与训练策略。代码、数据与在线排行榜见https://m-a-p.ai/OmniBench。

原文摘要 · Abstract (English)

Recent advancements in multimodal large language models (MLLMs) have aimed to integrate and interpret data across diverse modalities. However, the capacity of these models to concurrently process and reason about multiple modalities remains underexplored, partly due to the lack of comprehensive modality-wise benchmarks. We introduce OmniBench, a novel benchmark designed to rigorously evaluate models' ability to recognize, interpret, and reason across visual, acoustic, and textual inputs simultaneously. We define language models capable of such tri-modal processing as the omni-language models (OLMs). OmniBench is distinguished by high-quality human annotations, ensuring that accurate responses require integrated understanding and reasoning across all three modalities. Our main findings reveal that: i) open-source OLMs exhibit critical limitations in instruction-following and reasoning capabilities within tri-modal contexts; and ii) most baselines models perform poorly (below 50% accuracy) even when provided with alternative textual representations of images or/and audio. These results suggest that the ability to construct a consistent context from text, image, and audio is often overlooked in existing MLLM training paradigms. To address this gap, we curate an instruction tuning dataset of 84.5K training samples, OmniInstruct, for training OLMs to adapt to tri-modal contexts. We advocate for future research to focus on developing more robust tri-modal integration techniques and training strategies to enhance OLMs. Codes, data and live leaderboard could be found at https://m-a-p.ai/OmniBench.

多模态三模态基准测试语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。