arXiv:2508.04043cs.CV2025-08被引 12

首个面向真实人机交互场景的视觉变换推理基准,评测模型对动态变化的理解能力。

VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning

  • 基于第一视角操作视频构建,覆盖12类任务与3种推理维度
  • 模型在动态多步推理中表现不佳,尤其在中间状态识别上存在明显短板
  • 适合研究视觉-语言模型在因果推理与时序理解方面的能力

视觉变换推理(VTR)是智能体理解动态场景、建模因果关系并预测未来状态的关键认知能力,为高级智能系统奠定基础。然而现有基准存在仿真到现实的差距、任务复杂度有限及推理覆盖不全等问题,难以支撑真实场景应用。为此,我们提出VisualTrans,首个专为真实人机交互场景设计的VTR综合基准。该基准包含12种语义多样化的操作任务,通过6种明确的子任务类型,系统评估空间、流程与量化三个核心推理维度。数据集涵盖472组高质量问答对,形式包括单选题、开放式计数和目标枚举。我们构建了可扩展的数据生成流程,依托第一视角操作视频,整合任务选择、图像对提取、基于大模型的自动化元数据标注及结构化问题生成,并经人工验证确保数据质量与可解释性。对多种先进视觉-语言模型的评估显示,其在静态空间任务中表现良好,但在动态多步推理场景中暴露显著缺陷,尤其在中间状态识别与变换序列规划方面。这些发现揭示了当前模型在时间建模与因果推理方面的根本不足,为未来研究指明方向。数据与代码已公开于https://github.com/WangYipu2002/VisualTrans。

原文摘要 · Abstract (English)

Visual transformation reasoning (VTR) is a vital cognitive capability that empowers intelligent agents to understand dynamic scenes, model causal relationships, and predict future states, and thereby guiding actions and laying the foundation for advanced intelligent systems. However, existing benchmarks suffer from a sim-to-real gap, limited task complexity, and incomplete reasoning coverage, limiting their practical use in real-world scenarios. To address these limitations, we introduce VisualTrans, the first comprehensive benchmark specifically designed for VTR in real-world human-object interaction scenarios. VisualTrans encompasses 12 semantically diverse manipulation tasks and systematically evaluates three essential reasoning dimensions - spatial, procedural, and quantitative - through 6 well-defined subtask types. The benchmark features 472 high-quality question-answer pairs in various formats, including multiple-choice, open-ended counting, and target enumeration. We introduce a scalable data construction pipeline built upon first-person manipulation videos, which integrates task selection, image pair extraction, automated metadata annotation with large multimodal models, and structured question generation. Human verification ensures the final benchmark is both high-quality and interpretable. Evaluations of various state-of-the-art vision-language models show strong performance in static spatial tasks. However, they reveal notable shortcomings in dynamic, multi-step reasoning scenarios, particularly in areas like intermediate state recognition and transformation sequence planning. These findings highlight fundamental weaknesses in temporal modeling and causal reasoning, providing clear directions for future research aimed at developing more capable and generalizable VTR systems. The dataset and code are available at https://github.com/WangYipu2002/VisualTrans.

视觉推理人机交互因果推理基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。