用高斯表示生成文本描述的6自由度位姿分布,提升大规模场景定位精度。
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation
- 基于扩散模型优化噪声位姿,结合文本编码器生成条件分布。
- 在五个大规模数据集上显著优于传统分布估计方法,定位更准确。
- 适合需要精准视觉语义定位的研究者,如自动驾驶场景理解。
在大规模3D场景中定位文本描述存在固有歧义,例如识别城市中的所有交通灯。为此,我们提出一种方法,根据文本描述生成相机位姿的概率分布,以支持对宽泛概念的鲁棒推理。该方法采用基于扩散的架构,将噪声6DoF相机位姿逐步修正至合理位置,条件信号来自预训练文本编码器。通过与预训练视觉-语言模型CLIP集成,建立文本描述与位姿分布间的强关联。利用3D高斯点云渲染候选位姿,通过视觉推理纠正错位样本,进一步提升定位精度。我们在五个大规模数据集上验证了该方法的优越性,相比标准分布估计方法表现更优。代码、数据集及更多信息将在项目页面公开。
原文摘要 · Abstract (English)
Localizing textual descriptions within large-scale 3D scenes presents inherent ambiguities, such as identifying all traffic lights in a city. Addressing this, we introduce a method to generate distributions of camera poses conditioned on textual descriptions, facilitating robust reasoning for broadly defined concepts. Our approach employs a diffusion-based architecture to refine noisy 6DoF camera poses towards plausible locations, with conditional signals derived from pre-trained text encoders. Integration with the pretrained Vision-Language Model, CLIP, establishes a strong linkage between text descriptions and pose distributions. Enhancement of localization accuracy is achieved by rendering candidate poses using 3D Gaussian splatting, which corrects misaligned samples through visual reasoning. We validate our method's superiority by comparing it against standard distribution estimation methods across five large-scale datasets, demonstrating consistent outperformance. Code, datasets and more information will be publicly available at our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。