用混合柯西模型提升文本描述与点云定位的匹配精度
CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based Framework
- 引入柯西混合模型建模文本与点云间的不确定语义关系
- 在KITTI360Pose数据集上达到最新最优性能,定位误差更低
- 适合处理部分描述场景,如车辆接驾等实际应用
基于语言描述的点云定位旨在大都市环境中通过文本确定三维位置,潜在应用于车辆接驳或货物配送。理想情况下,文本应完整描述目标位置周围的物体,但实际中用户通常仅描述最显著且临近的局部环境,构成‘部分相关’挑战。为此,本文提出CMMLoc——一种基于柯西混合模型(CMM)的不确定性感知框架,用于文本到点云的定位。通过在跨模态交互中引入CMM先验,建模文本与点云间的不确定语义关联;设计空间融合机制,实现不同接收域的3D物体自适应聚合;提出方位角融合模块与模态预对齐策略,增强物体间空间关系捕捉能力,使3D对象更贴近文本语义。大量实验验证,CMMLoc在KITTI360Pose数据集上表现优异,优于现有方法,达到当前最优水平。代码已开源。
原文摘要 · Abstract (English)
The goal of point cloud localization based on linguistic description is to identify a 3D position using textual description in large urban environments, which has potential applications in various fields, such as determining the location for vehicle pickup or goods delivery. Ideally, for a textual description and its corresponding 3D location, the objects around the 3D location should be fully described in the text description. However, in practical scenarios, e.g., vehicle pickup, passengers usually describe only the part of the most significant and nearby surroundings instead of the entire environment. In response to this $\textbf{partially relevant}$ challenge, we propose $\textbf{CMMLoc}$, an uncertainty-aware $\textbf{C}$auchy-$\textbf{M}$ixture-$\textbf{M}$odel ($\textbf{CMM}$) based framework for text-to-point-cloud $\textbf{Loc}$alization. To model the uncertain semantic relations between text and point cloud, we integrate CMM constraints as a prior during the interaction between the two modalities. We further design a spatial consolidation scheme to enable adaptive aggregation of different 3D objects with varying receptive fields. To achieve precise localization, we propose a cardinal direction integration module alongside a modality pre-alignment strategy, helping capture the spatial relationships among objects and bringing the 3D objects closer to the text modality. Comprehensive experiments validate that CMMLoc outperforms existing methods, achieving state-of-the-art results on the KITTI360Pose dataset. Codes are available in this GitHub repository https://github.com/kevin301342/CMMLoc.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。