解决自动驾驶中3D视觉定位的上下文理解难题
LidaRefer: Context-aware Outdoor 3D Visual Grounding for Autonomous Driving
- 通过对象中心特征选择聚焦关键视觉信息
- 在Talk2Car-3D上达到当前最优性能
- 适合自动驾驶场景下的多物体定位研究
3D视觉定位旨在根据自然语言描述定位3D场景中的物体或区域。尽管室内3D视觉定位已取得进展,但室外场景因存在大量背景点、前景信息稀少,且多数数据集缺乏参照性非目标物体的空间标注,导致跨模态对齐与上下文理解困难。为此,本文提出LidaRefer框架,采用基于对象的特征选择策略,减少计算开销并聚焦语义相关特征;其基于Transformer的编码器-解码器结构能实现精细的跨模态对齐与全局上下文建模。此外,提出判别-支持协同定位(DiSCo)监督策略,显式建模目标物、上下文物与模糊物间的空间关系以提升定位精度。为避免人工标注,引入伪标签方法自动获取参照性非目标物体的3D定位标签。LidaRefer在Talk2Car-3D数据集上多种评估设置下均达到当前最优表现。
原文摘要 · Abstract (English)
3D visual grounding (VG) aims to locate objects or regions within 3D scenes guided by natural language descriptions. While indoor 3D VG has advanced, outdoor 3D VG remains underexplored due to two challenges: (1) large-scale outdoor LiDAR scenes are dominated by background points and contain limited foreground information, making cross-modal alignment and contextual understanding more difficult; and (2) most outdoor datasets lack spatial annotations for referential non-target objects, which hinders explicit learning of referential context. To this end, we propose LidaRefer, a context-aware 3D VG framework for outdoor scenes. LidaRefer incorporates an object-centric feature selection strategy to focus on semantically relevant visual features while reducing computational overhead. Then, its transformer-based encoder-decoder architecture excels at establishing fine-grained cross-modal alignment between refined visual features and word-level text features, and capturing comprehensive global context. Additionally, we present Discriminative-Supportive Collaborative localization (DiSCo), a novel supervision strategy that explicitly models spatial relationships between target, contextual, and ambiguous objects for accurate target identification. To enable this without manual labeling, we introduce a pseudo-labeling approach that retrieves 3D localization labels for referential non-target objects. LidaRefer achieves state-of-the-art performance on Talk2Car-3D dataset under various evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。