让视觉思维真正有用:通过遮蔽视角引导模型依赖生成的思考图
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning

- 训练时用视图遮蔽(VDrop)强制模型依赖生成的思考图像
- 全景式思考图在五个真实场景测试中表现最佳,泛化能力最强
- 该方法适合需要跨视角空间推理的多模态系统开发者
跨视图空间推理仍是视觉语言模型的薄弱环节:模型常依赖语言推理而忽略精细几何信息。视觉思维通过生成中间思考图像来解决此问题,但现有研究发现模型常忽视这些视觉线索。为此,本文探讨如何使视觉思维有效,以及何种形式最有效。研究聚焦于原生支持图像-文本交替生成的统一多模态模型(UMMs)。针对第一个问题,提出训练时的视图遮蔽(VDrop)策略:隐藏输入视图的部分内容,使其不可见于答案区间,但仍对思考图像可见,从而激励模型使用思考图像进行回答。一旦确认思考图像被利用,进一步比较三种思考图像形式:自上而下、全景式与点匹配渲染。在合成场景训练、五个真实世界域外基准测试下,全景式视觉思维结合VDrop是唯一兼具可学习性与信息量的配置,实现了最佳的域外泛化性能。
原文摘要 · Abstract (English)
Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they often reason in language and lose the fine-grained geometry needed for the task. Thinking with images aims to address this by generating an intermediate thinking image, but recent work shows that models often ignore the visual evidence in these traces. We therefore ask how to make visual thinking matter, and what kind of visual thinking works best. We study these questions in unified multimodal models (UMMs), which natively support interleaved image-text generation. For the first question, we propose View Dropout (VDrop), a training-time intervention that hides parts of one input view from the answer span while keeping them visible to the thinking-image tokens. This encourages the model to use the thinking image when answering, instead of relying only on the input views. Once the thinking image is used for answer prediction, we study which type of visual thinking is most effective. We frame this as a learnability-informativeness tradeoff and compare three thinking-image variants: top-down, panoramic, and point-matching renderings. Trained on synthetic scenes and evaluated on five real-world out-of-domain benchmarks, panoramic visual thinking with VDrop is the only configuration that is both informative and learnable, and it achieves the best out-of-domain generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。