让大模型通过生成新视角来增强空间推理能力
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence

- 在推理中动态生成新视角,解决单一视角下的理解局限
- 提升空间推理准确率1.3至3.9个百分点,尤其在依赖视角的任务上效果显著
- 适合研究多视图推理、视觉智能与生成模型融合的学者
当前大型多模态模型(LMMs)在需要视角依赖理解的空间推理任务中表现不佳,主要因其仅依赖单一静态观测。我们提出「思维新视角」(TwNV)范式,将生成式新视角合成融入推理流程:一个推理型LMM识别空间模糊性,指令一个绘图模块生成替代视角,并利用新增证据重新审视场景。系统实验回答三个问题:(1) 数值化相机位姿指令比自由语言指令更可靠地控制视角;(2) 合成视角质量与下游空间推理准确性高度相关;(3) 推理时迭代多轮视角优化进一步提升性能,呼应语言推理中的近期扩展趋势。在四个空间子任务类别和四种LMM架构(含闭源与开源)上,TwNV一致提升准确率1.3至3.9个百分点,对视角敏感子任务提升最大。结果证明新视角生成是推动LMM空间智能发展的实用手段。
原文摘要 · Abstract (English)
Current Large Multimodal Models (LMMs) struggle with spatial reasoning tasks requiring viewpoint-dependent understanding, largely because they are confined to a single, static observation. We propose Thinking with Novel Views (TwNV), a paradigm that integrates generative novel-view synthesis into the reasoning loop: a Reasoner LMM identifies spatial ambiguity, instructs a Painter to synthesize an alternative viewpoint, and re-examines the scene with the additional evidence. Through systematic experiments we address three research questions. (1) Instruction format: numerical camera-pose specifications yield more reliable view control than free-form language. (2) Generation fidelity: synthesized view quality is tightly coupled with downstream spatial accuracy. (3) Inference-time visual scaling: iterative multi-turn view refinement further improves performance, echoing recent scaling trends in language reasoning. Across four spatial subtask categories and four LMM architectures (both closed- and open-source), TwNV consistently improves accuracy by +1.3 to +3.9 pp, with the largest gains on viewpoint-sensitive subtasks. These results establish novel-view generation as a practical lever for advancing spatial intelligence of LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。