arXiv:2412.06292cs.CV2024-12ICCV被引 3

用大模型零样本检测3D形状关键点,无需标注数据

ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models

  • 利用多模态大模型的视觉语言知识进行点级推理
  • 在标准基准上性能媲美有监督方法,且无需3D标注
  • 适合需要快速适配新类别或无标注场景的研究者

我们提出一种新颖的零样本3D关键点检测方法。点级视觉推理极具挑战性,即使强大模型如DINO或CLIP也难以精确定位。传统3D关键点检测依赖大量标注数据和监督训练,限制了其可扩展性和对新类别/领域的适用性。本方法利用多模态大语言模型(MLLMs)中嵌入的丰富知识,首次证明:可用于训练近期MLLMs的像素级标注,可直接用于提取和命名3D模型上的显著关键点,无需任何真实标签或监督。实验表明,该方法在标准基准上表现与有监督方法相当,且训练过程不需任何3D关键点标注。结果凸显了将语言模型融入局部3D理解的潜力,为跨模态学习开辟新路径,并验证了MLLMs在3D计算机视觉任务中的有效性。

原文摘要 · Abstract (English)

We propose a novel zero-shot approach for keypoint detection on 3D shapes. Point-level reasoning on visual data is challenging as it requires precise localization capability, posing problems even for powerful models like DINO or CLIP. Traditional methods for 3D keypoint detection rely heavily on annotated 3D datasets and extensive supervised training, limiting their scalability and applicability to new categories or domains. In contrast, our method utilizes the rich knowledge embedded within Multi-Modal Large Language Models (MLLMs). Specifically, we demonstrate, for the first time, that pixel-level annotations used to train recent MLLMs can be exploited for both extracting and naming salient keypoints on 3D models without any ground truth labels or supervision. Experimental evaluations demonstrate that our approach achieves competitive performance on standard benchmarks compared to supervised methods, despite not requiring any 3D keypoint annotations during training. Our results highlight the potential of integrating language models for localized 3D shape understanding. This work opens new avenues for cross-modal learning and underscores the effectiveness of MLLMs in contributing to 3D computer vision challenges.

3D关键点零样本大模型点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。