新基准DiCoBench测试模型在高分辨率图像中捕捉细微差异与共性特征的能力。
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues

- 设计双轨多图高分辨率评测,聚焦隐式视觉线索的细粒度感知。
- 18个主流多模态大模型在该基准上表现远低于人类准确率(98.3%)。
- 适合研究高分辨率图像理解、细粒度感知与自主视觉推理的学者。
近年来,多模态大语言模型(MLLMs)在细粒度感知方面表现出色。然而,现有基准大多依赖显式文本提示或低分辨率输入,无法评估模型在高分辨率图像中自主感知隐式视觉线索的能力。为此,我们提出DiCoBench,一个全面的、多图像高分辨率基准,用于跨图像细粒度感知。DiCoBench包含765个精心筛选的样本,分为两个递进赛道:差异性视觉线索与共性视觉线索,涵盖8种感知任务。通过将评测设计为多选题形式并使用接近2K分辨率的图像,我们消除了评估指标偏差,对当前最先进的MLLMs提出了巨大挑战。对18种不同MLLMs的广泛评估显示,其性能与人类准确率(98.3%)存在显著差距,顶级模型在微尺度细节捕捉上仍明显不足。我们认为DiCoBench将成为推动未来自主高分辨率多图像感知研究的重要测试平台。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive fine-grained perception capabilities. However, existing benchmarks predominantly rely on explicit textual cues or low-resolution inputs, failing to evaluate a model's ability to autonomously perceive implicit visual cues in high-resolution. To bridge this gap, we introduce DiCoBench, a comprehensive, multi-image high-resolution benchmark designed for cross-image fine-grained perception. DiCoBench consists of 765 meticulously curated samples categorized into two progressive tracks: Differential Visual Cues and Commonality Visual Cues, covering 8 distinct perception tasks. By formulating the benchmark as a multiple-choice question task and utilizing high-resolution imagery (approaching 2K), we eliminate evaluation metric bias and pose a substantial challenge to current state-of-the-art MLLMs. Our extensive evaluation of 18 diverse MLLMs reveals a striking performance gap compared to human accuracy (98.3\%), with top-performing models struggling significantly with micro-scale detail capture. We believe DiCoBench will serve as a challenging testbed to drive future research in autonomous, high-resolution multi-image perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。