arXiv:2505.11838cs.CV2025-05被引 8

构建首个统一视觉推理基准,支持多模态输出与复杂推理链。

RVTBench: A Benchmark for Visual Reasoning Tasks

  • 用数字孪生作为感知与文本查询间的结构化中介,生成高质量推理数据。
  • 包含3896个查询、超120万词元,覆盖4类任务与4级难度。
  • 适合研究视频理解、多模态推理与零样本泛化方向的学者使用。

视觉推理指模型通过多步推理回应隐式文本查询来理解视觉输入的能力,但深度学习模型仍面临挑战,主要源于缺乏相关基准。以往研究多集中于推理分割任务,即根据隐式文本查询分割物体。本文提出推理视觉任务(RVT),一种统一框架,将传统视频推理分割拓展至多样化的视觉语言推理问题,支持边界框、自然语言描述及问答对等多种输出格式。针对现有基准依赖大语言模型(LLM)导致空间-时间关系与多步推理链刻画不足的问题,我们提出一种新型自动化RVT基准构建流程,利用数字孪生(DT)作为感知与隐式文本查询生成之间的结构化中间体。基于此方法,我们构建了RVTBench,一个包含3896个查询、超过120万词元的基准,涵盖4种RVT类型(分割、定位、VQA、摘要)、3类推理类别(语义、空间、时间)和4个递增难度级别,源自200段视频序列。最后,我们提出RVTagent,一种无需任务微调即可实现跨多种RVT的零样本泛化代理框架。

原文摘要 · Abstract (English)

Visual reasoning, the capability to interpret visual input in response to implicit text query through multi-step reasoning, remains a challenge for deep learning models due to the lack of relevant benchmarks. Previous work in visual reasoning has primarily focused on reasoning segmentation, where models aim to segment objects based on implicit text queries. This paper introduces reasoning visual tasks (RVTs), a unified formulation that extends beyond traditional video reasoning segmentation to a diverse family of visual language reasoning problems, which can therefore accommodate multiple output formats including bounding boxes, natural language descriptions, and question-answer pairs. Correspondingly, we identify the limitations in current benchmark construction methods that rely solely on large language models (LLMs), which inadequately capture complex spatial-temporal relationships and multi-step reasoning chains in video due to their reliance on token representation, resulting in benchmarks with artificially limited reasoning complexity. To address this limitation, we propose a novel automated RVT benchmark construction pipeline that leverages digital twin (DT) representations as structured intermediaries between perception and the generation of implicit text queries. Based on this method, we construct RVTBench, a RVT benchmark containing 3,896 queries of over 1.2 million tokens across four types of RVT (segmentation, grounding, VQA and summary), three reasoning categories (semantic, spatial, and temporal), and four increasing difficulty levels, derived from 200 video sequences. Finally, we propose RVTagent, an agent framework for RVT that allows for zero-shot generalization across various types of RVT without task-specific fine-tuning.

视觉推理多模态数字孪生基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。