首个统一多模态计数评测基准,解决视觉、文本、音频计数评估标准不一问题。
UNICBench: UNIfied Counting Benchmark for MLLM
- 构建涵盖图像、文档、音频的统一计数数据集,支持多层级能力评估。
- 45个主流多模态大模型测试显示,基础任务表现好但复杂推理仍有明显差距。
- 提供标准化评测协议与开源工具包,适合研究者提升模型计数能力。
计数是多模态大语言模型的核心能力,但目前缺乏跨视觉、文本和音频的统一评测数据集。我们提出UNICBench,一个统一的多模态、多层次计数评测基准与工具包,具备精确的标注、确定性数值解析和分层报告机制。数据集包含5,300张图像(5,508个问答)、872份文档(5,888个问答)和2,069段音频(2,905个问答),并标注了三级能力分类和难度标签。在固定划分、提示和种子的标准协议下,对45个顶尖多模态大模型进行跨模态评估。结果表明,模型在基础计数任务上表现良好,但在推理和最困难分区存在显著差距,揭示长尾错误和巨大提升空间。UNICBench为评估提供严谨可比的基准,并开放工具包以推动进展。
原文摘要 · Abstract (English)
Counting is a core capability for multimodal large language models (MLLMs), yet there is no unified counting dataset to rigorously evaluate this ability across image, text, and audio. We present UNICBench, a unified multimodal, multi level counting benchmark and evaluation toolkit with accurate ground truth, deterministic numeric parsing, and stratified reporting. The corpus comprises 5,300 images (5,508 QA), 872 documents (5,888 QA), and 2,069 audio clips (2,905 QA), annotated with a three level capability taxonomy and difficulty tags. Under a standardized protocol with fixed splits/prompts/seeds and modality specific matching rules, we evaluate 45 state-of-the-art MLLMs across modalities. Results show strong performance on some basic counting tasks but significant gaps on reasoning and the hardest partitions, highlighting long-tail errors and substantial headroom for improving general counting. UNICBench offers a rigorous and comparable basis for measurement and a public toolkit to accelerate progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。