用几何确定性自动生成3D空间推理标签,解决模型自我纠错难题。
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments

- 基于点云和相机位姿计算真实答案,无需模型参与
- 在9个基准上3B/7B模型均达最高平均分,且不损害通用视觉理解
- 自动聚焦弱项任务,实现无需人工设计的动态训练课程
三维场景中的空间推理是具身智能的核心能力,但持续改进受限于几何标注成本。自演化范式虽有潜力,却因依赖模型共识生成伪标签,反而强化了模型自身的几何错误。我们发现三维空间推理的独特性质:真实答案是底层几何的确定性结果,可从点云和相机位姿精确计算,无需模型介入。基于此,提出SpatialEvo框架,核心为确定性几何环境(DGE)。DGE将16类空间推理任务形式化为显式几何验证规则,将未标注的3D场景转化为零噪声交互式原语,取代模型共识,提供客观物理反馈。一个共享参数策略在DGE约束下同时进化提问者与求解者角色:提问者生成基于场景观测的物理有效问题,求解者对经DGE验证的真实答案进行精确推导。任务自适应调度器自主聚焦模型最薄弱类别,生成无需人工设计的动态课程。在九个基准上的实验表明,SpatialEvo在3B和7B规模下均取得最高平均分,空间推理性能持续提升,通用视觉理解无退化。
原文摘要 · Abstract (English)
Spatial reasoning over three-dimensional scenes is a core capability for embodied intelligence, yet continuous model improvement remains bottlenecked by the cost of geometric annotation. The self-evolving paradigm offers a promising path, but its reliance on model consensus to construct pseudo-labels causes training to reinforce rather than correct the model's own geometric errors. We identify a property unique to 3D spatial reasoning that circumvents this limitation: ground truth is a deterministic consequence of the underlying geometry, computable exactly from point clouds and camera poses without any model involvement. Building on this insight, we present SpatialEvo, a self-evolving framework for 3D spatial reasoning, centered on the Deterministic Geometric Environment (DGE). The DGE formalizes 16 spatial reasoning task categories under explicit geometric validation rules and converts unannotated 3D scenes into zero-noise interactive oracles, replacing model consensus with objective physical feedback. A single shared-parameter policy co-evolves across questioner and solver roles under DGE constraints: the questioner generates physically valid spatial questions grounded in scene observations, while the solver derives precise answers against DGE-verified ground truth. A task-adaptive scheduler endogenously concentrates training on the model's weakest categories, producing a dynamic curriculum without manual design. Experiments across nine benchmarks demonstrate that SpatialEvo achieves the highest average score at both 3B and 7B scales, with consistent gains on spatial reasoning benchmarks and no degradation on general visual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。