统一空间感知与推理,用生成式模型一次搞定3D重建、对应关系和空间理解。
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation

- 将3D重建、对应匹配和空间推理都当作生成任务处理。
- 在多个基准上表现优异,单一框架完成多种空间任务。
- 适合研究多模态生成与空间认知融合的学者参考。
从视觉观察中进行空间感知与推理需恢复几何结构、建立对应关系并理解空间关系。现有方法通常采用专用架构或外部几何模块分别处理这些能力,限制了对同一物理场景互补表征的知识迁移。我们提出SPARGen,一个统一的多模态框架,将3D重建、密集对应关系和空间推理均视为指令条件生成任务。SPARGen将紧凑的结构化与语言输出序列化为标记序列,同时以图像对齐形式生成密集几何场,使空间监督能共同塑造共享表示。在3D重建、对应关系和空间推理的多个基准测试中,SPARGen在单一原生多模态生成框架内实现了具有竞争力的表现。
原文摘要 · Abstract (English)
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。