arXiv:2607.12800cs.CV2026-07

让模型从纯视觉中学会复杂推理与物理规划。

UniVR: Thinking in Visual Space for Unified Visual Reasoning

论文配图:UniVR: Thinking in Visual Space for Unified Visual Reasoning
图 1 · 摘自论文原文
  • 用全局与步级奖励协同优化视觉推理过程。
  • 在VR-X上性能提升25%,超越多模态基准。
  • 适合研究视觉智能与通用推理的学者。

从原始视觉数据中直接学习广泛世界知识是智能的核心能力。我们提出UniVR,首次实现从纯视觉演示中同步学习复杂推理、精细物理动态和长期规划。其核心是VR-GRPO强化学习范式,通过互补的全局与步级奖励,确保推理过程逻辑连贯与物理一致,无需任务特定启发式或图文对。为训练与评估UniVR,我们构建了涵盖16个来源的大型基准VR-X,覆盖长时程操作、空间谜题与物理推理,是首个在纯视觉协议下评估这些异构能力的综合套件。惊人的是,UniVR在VR-X上最高提升25%,其卓越视觉推理能力也显著改善了多种多模态理解基准表现。这些发现凸显了视觉空间内推理的巨大潜力,所有代码、数据与模型均已开源。

原文摘要 · Abstract (English)

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.

视觉推理强化学习物理建模通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。