通过密集奖励提升多视角3D推理的连贯性与视点选择能力。
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views

- 构建全局地图并规划问题引导的视点轨迹,实现分步推理。
- 在三个数据集上超越强基线,提升跨视角一致性与答案准确率。
- 适合研究多视角3D理解、视觉语言模型推理的学者使用。
多视角3D视觉问答(MV3D-VQA)需将局部观测整合为一致的3D场景表示,并选出有助于多步空间推理的视角。现有多模态大模型通常仅以稀疏的答案级监督训练,导致跨视角推理不一致且视点选择脆弱。本文提出DR-MV3D(密集奖励用于多视角3D-VQA),一个基于地图的训练框架,提供可验证的密集奖励以监督推理过程。该方法将任务分解为:(i) 构建以环境为中心的全局地图,(ii) 根据问题生成视点轨迹,(iii) 以自身为中心定位进行答案预测。为使中间步骤无需人工标注即可学习,引入两项奖励:全局一致性奖励,将预测地图与来自冻结3D视觉基础模型(如VGGT + SAM3)的几何一致伪目标对齐;局部轨迹奖励,监督视点的有序选择。通过轨迹级策略优化(GRPO)联合优化整个流程。在MindCube、VSI-Bench和BLINK(MV)上的实验表明,DR-MV3D持续优于多个强基线,验证了过程级密集监督在多视角3D推理中的有效性。
原文摘要 · Abstract (English)
Multi-view 3D Visual Question Answering (MV3D-VQA) requires integrating partial observations into a coherent 3D scene representation and selecting informative viewpoints for multi-step spatial reasoning. However, current multimodal LLMs are typically trained with sparse, answer-level supervision, which often yields inconsistent cross-view reasoning and brittle view selection. We present DR-MV3D (Dense Reward for MV3D-VQA), a map-grounded learning framework that provides dense, verifiable rewards to supervise the reasoning process. Our approach decomposes MV3D-VQA into (i) allocentric global map construction, (ii) question-conditioned view-trajectory planning, and (iii) egocentric grounding for answer prediction. To make intermediate steps learnable without manual annotations, we introduce two rewards: a global consistency reward that aligns the predicted map with geometry-consistent pseudo targets from frozen 3D vision foundation models (e.g., VGGT + SAM3), and a local trajectory reward that supervises ordered viewpoint selection. We optimize the full pipeline with trajectory-level policy optimization (GRPO). Experiments on MindCube, VSI-Bench, and BLINK (MV) show that DR-MV3D consistently improves over strong multi-image baselines, supporting the effectiveness of process-level dense supervision for multi-view 3D reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。