arXiv:2603.06024cs.CLcs.CV2026-03被引 1

让视觉语言模型学会多视角空间推理,提升复杂场景理解能力

ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoning

  • 分两阶段处理:先推断视角间空间关系,再基于中间工作空间回答问题
  • 在MMSI-Bench上准确率提升5.3%,对需跨视图对齐的任务增益更明显
  • 适合需要多视角推理的场景,如机器人导航、三维重建等应用

多视角空间推理仍是当前视觉语言模型的难点。即使提供多个视角,模型常忽视跨视图关系,依赖单图捷径,导致在视角变换和遮挡敏感任务中表现脆弱。我们提出ViewFusion,一种两阶段框架,明确分离跨视图空间预对齐与问答过程。第一阶段,模型进行主动空间预思考,推断各视角间的空间关系与变换,构建超越简单重述的中间工作空间。第二阶段,模型基于该工作空间进行问题驱动推理,生成最终答案。采用合成推理监督训练后,通过GRPO强化学习优化,提升答案正确性并稳定两阶段生成行为。在MMSI-Bench上,ViewFusion相比Qwen3-VL-4B-Instruct准确率提升5.3%,尤其在需真实跨视图对齐的任务中表现突出。

原文摘要 · Abstract (English)

Multi-view spatial reasoning remains difficult for current vision-language models. Even when multiple viewpoints are available, models often underutilize cross-view relations and instead rely on single-image shortcuts, leading to fragile performance on viewpoint transformation and occlusion-sensitive cases. We present ViewFusion, a two-stage framework that explicitly separates cross-view spatial pre-alignment from question answering. In the first stage, the model performs deliberate spatial pre-thinking to infer viewpoint relations and spatial transformations across views, forming an intermediate workspace that goes beyond a simple re-description. In the second stage, the model conducts question-driven reasoning conditioned on this workspace to produce the final prediction. We train ViewFusion with synthetic reasoning supervision followed by reinforcement learning using GRPO, which improves answer correctness while stabilizing the intended two-stage generation behavior. On MMSI-Bench, ViewFusion improves accuracy by 5.3\% over Qwen3-VL-4B-Instruct, with the largest gains on examples that require genuine cross-view alignment.

多视角推理空间思维链视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。