构建统一基准,揭示单模与全模模型能力的关联规律
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- 设计统一评估框架,覆盖44类任务和5种模态组合
- 发现强模型全模能力呈协同提升,弱模型受瓶颈制约
- 含中文场景数据,支持自动评分,适配模型研发与评测
多模态大语言模型正从单模理解向融合视觉、音频与语言的全模模型演进。然而,单模与全模能力之间的关联尚不明确,亟需全面评估以推动全模模型智能发展。本文提出首个统一的全模基准UNO-Bench,可统一评估单模与全模能力,涵盖44类任务及5种模态组合。包含1250个人工标注的全模样本(跨模态可解率达98%),以及2480个增强的单模样本。人工数据贴合真实场景,尤其适用于中文环境;自动生成数据提速90%,在18个公开基准上保持98%一致性。除传统选择题外,新增多步开放式问题以评估复杂推理。引入通用评分模型,支持6类问题自动化评估,准确率达95%。实验揭示全模与单模间存在组合规律:弱模型中全模能力成瓶颈,而强模型则呈现协同增益。
原文摘要 · Abstract (English)
Multimodal Large Languages models have been progressing from uni-modal understanding toward unifying visual, audio and language modalities, collectively termed omni models. However, the correlation between uni-modal and omni-modal remains unclear, which requires comprehensive evaluation to drive omni model's intelligence evolution. In this work, we introduce a novel, high-quality, and UNified Omni model benchmark, UNO-Bench. This benchmark is designed to effectively evaluate both UNi-modal and Omni-modal capabilities under a unified ability taxonomy, spanning 44 task types and 5 modality combinations. It includes 1250 human curated samples for omni-modal with 98% cross-modality solvability, and 2480 enhanced uni-modal samples. The human-generated dataset is well-suited to real-world scenarios, particularly within the Chinese context, whereas the automatically compressed dataset offers a 90% increase in speed and maintains 98% consistency across 18 public benchmarks. In addition to traditional multi-choice questions, we propose an innovative multi-step open-ended question format to assess complex reasoning. A general scoring model is incorporated, supporting 6 question types for automated evaluation with 95% accuracy. Experimental result shows the Compositional Law between omni-modal and uni-modal performance and the omni-modal capability manifests as a bottleneck effect on weak models, while exhibiting synergistic promotion on strong models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。