arXiv:2605.18162cs.CVcs.AI2026-05

让视觉语言模型自动提升空间推理能力,应对变换后答案不一致问题。

Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency

论文配图:Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
图 1 · 摘自论文原文
  • 通过几何与语言双重操作,强制模型在原图和变换图上保持逻辑一致。
  • 在多个视频与空间推理数据集上,显著优于强基线模型。
  • 无需大量数据,可作为轻量级后训练模块适配现有模型。

视觉语言模型(VLMs)取得显著进展,但其空间推理仍脆弱:即使对原始输入回答正确,面对具有可预测答案映射的变换输入仍会失败,暴露出实例级正确性与鲁棒空间推理之间的差距。为此,我们提出空间对齐的几何进化框架SAGE,通过几何与语言双重性操作强化VLM中的逻辑一致性。SAGE将双重一致性作为辅助奖励嵌入GRPO训练中,促使模型在原始与变换输入间生成逻辑自洽的答案。动态操作池持续探测不一致,促进挑战性操作,淘汰已掌握的操作,使训练聚焦于最具信息量的信号。SAGE具有模型无关性,相比以往GRPO方法更数据高效,可作为轻量级后训练阶段应用于任意现有VLM。在视频与空间推理基准上的实验表明,SAGE持续优于强基线,并增强对未见数据的泛化能力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have made striking progress, yet their spatial reasoning remains fragile: models that answer an original input correctly can still fail under paired transformations with predictable answer mappings, revealing a gap between instance-level correctness and robust spatial reasoning. To address this, we propose Spatial Alignment via Geometric Evolution (SAGE), a self-evolving framework that enforces logical consistency in VLMs through geometric and linguistic duality operations. SAGE incorporates duality consistency as an auxiliary reward within GRPO training, encouraging models to produce logically coherent answers across original and transformed inputs. A dynamic operation pool continuously probes for inconsistencies, promoting challenging operations and retiring mastered ones, so that training focuses on the most informative signals. SAGE is model-agnostic, data-efficient compared to prior GRPO methods, and can be applied as a lightweight post-training stage to any existing VLM. Experiments on video and spatial reasoning benchmarks demonstrate consistent improvements over strong baselines and enhanced generalization to unseen data.

空间推理视觉语言模型自进化逻辑一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。