从无姿态多视角图像中学习鲁棒的3D表示,提升空间智能感知能力。
Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
- 采用双掩码策略增强编码器对几何结构的感知能力。
- 通过粗到精的高斯点阵优化,减少外观与语义不一致现象。
- 基于相机位姿重投影校准,实现几何与语义的跨任务一致性。
稳健的3D表征学习构成了空间智能的感知基础,支持场景理解与具身AI中的下游任务。然而,直接从无姿态多视角图像中学习此类表征仍具挑战性。现有自监督方法虽尝试统一几何、外观与语义,但常面临几何引导弱、外观细节不足及几何与语义不一致等问题。本文提出UniSplat,一个前馈式框架,通过三个互补组件解决上述问题:首先,设计双掩码策略,在编码器与解码器均施加掩码,并将解码器掩码聚焦于几何丰富区域,迫使模型从不完整视觉线索中推断结构信息,获得即使在无姿态输入下也具备几何感知力的表征;其次,提出粗到精的高斯点阵策略,逐步细化辐射场,降低外观与语义间的不一致性;最后,引入位姿条件下的再校准机制,利用估计的相机参数将预测的3D点云与语义图重投影至图像平面,与对应RGB和语义预测对齐,确保跨任务一致性,有效缓解几何-语义错配问题。三者协同,生成对无姿态、稀疏视图输入鲁棒且可泛化的统一3D表征,为构建空间智能的感知基础提供支撑。
原文摘要 · Abstract (English)
Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images remains challenging. Recent self-supervised methods attempt to unify geometry, appearance, and semantics in a feed-forward manner, but they often suffer from weak geometry induction, limited appearance detail, and inconsistencies between geometry and semantics. We introduce UniSplat, a feed-forward framework designed to address these limitations through three complementary components. First, we propose a dual-masking strategy that strengthens geometry induction in the encoder. By masking both encoder and decoder tokens, and targeting decoder masks toward geometry-rich regions, the model is forced to infer structural information from incomplete visual cues, yielding geometry-aware representations even under unposed inputs. Second, we develop a coarse-to-fine Gaussian splatting strategy that reduces appearance-semantics inconsistencies by progressively refining the radiance field. Finally, to enforce geometric-semantic consistency, we introduce a pose-conditioned recalibration mechanism that interrelates the outputs of multiple heads by re-projecting predicted 3D point and semantic maps into the image plane using estimated camera parameters, and aligning them with corresponding RGB and semantic predictions to ensure cross-task consistency, thereby resolving geometry-semantic mismatches. Together, these components yield unified 3D representations that are robust to unposed, sparse-view inputs and generalize across diverse tasks, laying a perceptual foundation for spatial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。