arXiv:2608.14138cs.CVcs.AI2026-08

统一空间感知与推理,用生成式模型一次搞定3D重建、对应关系和空间理解。

SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

论文配图:SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
图 1 · 摘自论文原文
  • 将3D重建、对应匹配和空间推理都当作生成任务处理。
  • 在多个基准上表现优异,单一框架完成多种空间任务。
  • 适合研究多模态生成与空间认知融合的学者参考。

从视觉观察中进行空间感知与推理需恢复几何结构、建立对应关系并理解空间关系。现有方法通常采用专用架构或外部几何模块分别处理这些能力,限制了对同一物理场景互补表征的知识迁移。我们提出SPARGen,一个统一的多模态框架,将3D重建、密集对应关系和空间推理均视为指令条件生成任务。SPARGen将紧凑的结构化与语言输出序列化为标记序列,同时以图像对齐形式生成密集几何场,使空间监督能共同塑造共享表示。在3D重建、对应关系和空间推理的多个基准测试中,SPARGen在单一原生多模态生成框架内实现了具有竞争力的表现。

原文摘要 · Abstract (English)

Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.

空间推理多模态生成3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。