arXiv:2507.12336cs.CV2025-07中稿 · CVPR

仅用一张图就能精准定位3D关键点,无需人工标注或多视角数据。

Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

  • 利用预训练多视角扩散模型生成多视图图像作为几何监督信号。
  • 在多个数据集上实现高精度3D关键点预测,跨域泛化能力强。
  • 可直接操控扩散模型生成的3D物体,适合3D内容创作与逆向设计。

现有3D关键点估计方法通常依赖人工标注或多视角校准图像,获取成本高。本文提出KeyDiff3D框架,仅需单张图像即可准确预测3D关键点,避免了昂贵的数据采集。该方法借助预训练多视角扩散模型中蕴含的强大几何先验,通过扩散模型从单图生成多视图图像,作为监督信号提供3D几何线索。我们还设计了一个3D特征提取器,将扩散特征中的隐式3D先验转换为显式的3D特征体。实验结果表明,该方法在Human3.6M、CUB-200-2011、Stanford Dogs等多样数据集及真实场景和域外输入上均表现出色,不仅关键点估计精度高,且具备操控扩散模型生成3D物体的能力。

原文摘要 · Abstract (English)

Most existing 3D keypoint estimation methods rely on manual annotations or calibrated multi-view images, both of which are expensive to collect. This paper introduces KeyDiff3D, a framework that can accurately predict 3D keypoints from a single image, thus eliminating the need for such expensive data acquisitions. To achieve this, we leverage powerful geometric priors embedded in a pretrained multi-view diffusion model. In our framework, the diffusion model generates multi-view images from a single image, serving as supervision signals to provide 3D geometric cues to our model. We also introduce a 3D feature extractor that transforms implicit 3D priors embedded in the diffusion features into explicit 3D feature volumes. Beyond accurate keypoint estimation, we further introduce a pipeline that enables manipulation of 3D objects generated by the diffusion model. Experimental results on diverse datasets, including Human3.6M, CUB-200-2011, Stanford Dogs, and several in-the-wild and out-of-domain inputs, highlight the effectiveness of our method in terms of accuracy, generalization, and its ability to enable manipulation of 3D objects generated by the diffusion model from a single image.

3D关键点单目图像扩散模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。