arXiv:2602.21992cs.CV2026-02被引 6

用强化学习提升视觉语言模型对全景图的3D空间理解能力

PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement Learning

  • 构建合成3D环境的全景问答基准PanoEnv,含14.8K题
  • 提出基于GRPO的强化学习框架,开环问答准确率提升至14.83%
  • 两阶段课程训练缓解遗忘,7B模型超越32B大模型表现

360度全景图像在虚拟现实、自动驾驶和机器人中广泛用于全局场景理解。然而,现有视觉语言模型(VLMs)在等距柱状投影(ERP)图像上因几何失真和缺乏3D监督而难以进行3D空间推理。我们提出PanoEnv,一个大规模的视觉问答基准,基于合成3D环境构建,包含14.8K个问题,涵盖相对位置、体积比较等五类,配有精确的3D标注(深度、分割、边界框)。对14个主流VLMs的评测显示,整体准确率仅49.34%,开放问答(OE)准确率仅为8.36%。为此,我们提出一种基于组相对策略优化(GRPO)的强化学习后训练框架,采用基于真实标注的奖励机制,融合距离容差、空间一致性等五种几何感知策略。通过两阶段课程训练:第一阶段在结构化任务(是非与多选)上训练,第二阶段在混合开放问答数据上微调以提升泛化能力。我们的7B模型达到新最优性能,整体准确率52.93%(+3.59%),开放问答准确率14.83%,同时保持结构化任务表现。其语义评估得分也领先(Q-Score 6.24,P-Score 5.95),超越32B模型。结果表明,PanoEnv-QA与课程式强化学习框架能有效赋予VLMs全景感知中的3D空间智能。

原文摘要 · Abstract (English)

360 panoramic images are increasingly used in virtual reality, autonomous driving, and robotics for holistic scene understanding. However, current Vision-Language Models (VLMs) struggle with 3D spatial reasoning on Equirectangular Projection (ERP) images due to geometric distortion and limited 3D supervision. We introduce PanoEnv, a large-scale VQA benchmark built from synthetic 3D environments, containing 14.8K questions across five categories (e.g., relative position, volume comparison) grounded in accurate 3D annotations including depth, segmentation, and bounding boxes. Benchmarking 14 state-of-the-art VLMs reveals limited 3D understanding, achieving only 49.34% overall accuracy and 8.36% on open-ended (OE) questions. To enhance 3D reasoning, we propose a reinforcement learning post-training framework based on Group Relative Policy Optimization (GRPO) with a ground-truth-guided reward that incorporates five geometry-aware strategies such as distance tolerance and spatial consistency. A two-stage curriculum further mitigates catastrophic forgetting: Stage 1 trains on structured tasks (true/false and multiple choice), and Stage 2 fine-tunes on mixed open-ended data to improve generalization. Our 7B model achieves new state-of-the-art performance, improving overall accuracy to 52.93% (+3.59%) and open-ended accuracy to 14.83% while maintaining structured-task performance. It also achieves top semantic evaluation scores (Q-Score 6.24, P-Score 5.95), surpassing 32B models. These results demonstrate that PanoEnv-QA and our curriculum-based RL framework effectively instill 3D spatial intelligence in VLMs for omnidirectional perception.

全景理解3D推理强化学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。