无需预设类别即可完成任意激光雷达物体的形状与识别。
Towards Learning to Complete Anything in Lidar
- 利用多模态时序数据挖掘物体形状与语义特征,实现零样本学习。
- 仅依赖部分补全信息,模型仍能推断出完整物体形状。
- 支持开放词汇识别,适用于标准场景补全基准测试。
我们提出CAL(Complete Anything in Lidar),用于真实场景下的基于激光雷达的形状补全。现有方法仅能对已有数据集中标注的封闭词汇物体进行补全与识别。而我们的零样本方法通过多模态传感器序列中的时序上下文,挖掘观测物体的形状与语义特征,并将其提炼为仅依赖激光雷达的实例级补全与识别模型。尽管仅挖掘了部分形状补全,但发现该模型可利用数据集中多个此类局部观察,推断出完整物体形状。实验表明,该模型可在标准语义与全景场景补全基准上进行提示推理,以(非可视)3D边界框定位物体,并识别超出固定类别词汇的物体。项目页面:https://research.nvidia.com/labs/dvl/projects/complete-anything-lidar
原文摘要 · Abstract (English)
We propose CAL (Complete Anything in Lidar) for Lidar-based shape-completion in-the-wild. This is closely related to Lidar-based semantic/panoptic scene completion. However, contemporary methods can only complete and recognize objects from a closed vocabulary labeled in existing Lidar datasets. Different to that, our zero-shot approach leverages the temporal context from multi-modal sensor sequences to mine object shapes and semantic features of observed objects. These are then distilled into a Lidar-only instance-level completion and recognition model. Although we only mine partial shape completions, we find that our distilled model learns to infer full object shapes from multiple such partial observations across the dataset. We show that our model can be prompted on standard benchmarks for Semantic and Panoptic Scene Completion, localize objects as (amodal) 3D bounding boxes, and recognize objects beyond fixed class vocabularies. Our project page is https://research.nvidia.com/labs/dvl/projects/complete-anything-lidar
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。