用几何方法提升文本定位精度,让机器人更懂人话。
Riemannian and Symplectic Geometry for Hierarchical Text-Driven Place Recognition
- 分层次对齐:实例、关系、全局三重匹配
- 在KITTI360上提升19%的定位准确率
- 适合做智能驾驶和配送的机器人定位
文本到点云定位使机器人通过自然语言理解空间位置,对自动驾驶和末端配送等场景至关重要。现有方法使用全局描述符进行相似性检索,存在信息损失且难以捕捉区分性场景结构。为此,我们提出SympLoc,一种粗粒度到细粒度的定位框架,包含多层级对齐机制。粗粒度阶段包含三个互补对齐层级:1)实例级对齐通过双曲空间中的黎曼自注意力建立点云中单个物体与文本提示间的直接对应;2)关系级对齐使用信息-辛关系编码器(ISRE),通过Fisher-Rao度量与哈密顿动力学重构关系特征,实现不确定性感知的几何一致性传播;3)全局级对齐通过谱流形变换(SMT)提取图谱分析得到的结构不变量,合成判别性全局描述符。该分层对齐策略逐步捕获从细粒度到粗粒度的场景语义,实现鲁棒跨模态检索。在KITTI360Pose数据集上的实验表明,SympLoc相较现有最先进方法在Top-1召回率@10m上提升19%。
原文摘要 · Abstract (English)
Text-to-point-cloud localization enables robots to understand spatial positions through natural language descriptions, which is crucial for human-robot collaboration in applications such as autonomous driving and last-mile delivery. However, existing methods employ pooled global descriptors for similarity retrieval, which suffer from severe information loss and fail to capture discriminative scene structures. To address these issues, we propose SympLoc, a novel coarse-to-fine localization framework with multi-level alignment in the coarse stage. Different from previous methods that rely solely on global descriptors, our coarse stage consists of three complementary alignment levels: 1) Instance-level alignment establishes direct correspondence between individual object instances in point clouds and textual hints through Riemannian self-attention in hyperbolic space; 2) Relation-level alignment explicitly models pairwise spatial relationships between objects using the Information-Symplectic Relation Encoder (ISRE), which reformulates relation features through Fisher-Rao metric and Hamiltonian dynamics for uncertainty-aware geometrically consistent propagation; 3) Global-level alignment synthesizes discriminative global descriptors via the Spectral Manifold Transform (SMT) that extracts structural invariants through graph spectral analysis. This hierarchical alignment strategy progressively captures fine-grained to coarse-grained scene semantics, enabling robust cross-modal retrieval. Extensive experiments on the KITTI360Pose dataset demonstrate that SympLoc achieves a 19% improvement in Top-1 recall@10m compared to existing state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。