arXiv:2601.08748cs.CVcs.AI2026-01被引 1

评测大模型在超高清图像上的多跳推理能力,填补高分辨率视觉理解空白。

UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images

  • 构建超高清图像多跳推理基准,覆盖人文与自然场景四类数据集
  • 图像分辨率达百兆至千兆像素,问题分三级评估推理深度
  • 提出基于智能体的推理框架,提升大模型处理超高清图像效率

近期多模态大语言模型在视觉-语言推理方面表现出色,但在超高清图像上的表现仍缺乏系统评估。现有视觉问答(VQA)基准多基于中等分辨率数据,视觉复杂度有限。为此,我们提出超高清推理基准(UR-Bench),旨在评估多模态大模型在极端视觉信息下的推理能力。UR-Bench包含人文场景与自然场景两大类别,涵盖四个子集,图像分辨率从数百兆像素到千兆像素不等,数据来源多样,空间结构各异。每个子集配备三层次问题,支持对模型推理能力的分级评估。我们进一步提出一种基于智能体的框架,通过调用外部视觉工具实现语言模型的逐步推理。引入语义抽象与检索工具,提升超高清图像处理效率。我们在端到端多模态大模型与自研框架下评估了前沿模型,验证了该框架的有效性。

原文摘要 · Abstract (English)

Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexplored. Existing visual question answering (VQA) benchmarks typically rely on medium-resolution data, offering limited visual complexity. To bridge this gap, we introduce Ultra-high-resolution Reasoning Benchmark (UR-Bench), a benchmark designed to evaluate the reasoning capabilities of MLLMs under extreme visual information. UR-Bench comprises two major categories, Humanistic Scenes and Natural Scenes, covering four subsets of ultra-high-resolution images with distinct spatial structures and data sources. Each subset contains images ranging from hundreds of megapixels to gigapixels, accompanied by questions organized into three levels, enabling evaluation of models' reasoning capabilities in ultra-high-resolution scenarios. We further propose an agent-based framework in which a language model performs reasoning by invoking external visual tools. In addition, we introduce Semantic Abstraction and Retrieval tools that enable more efficient processing of ultra-high-resolution images. We evaluate state-of-the-art models using both an end-to-end MLLMs and our agent-based framework, demonstrating the effectiveness of our framework.

多跳推理超高清图像视觉问答智能体框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。