arXiv:2504.21749cs.CV2025-04CVPR被引 11

无需标注数据,从视频中自监督学习常见物体的3D形状与外观模型。

Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature Space

  • 用神经特征替代颜色,通过可变形模板学习物体3D形态和外观。
  • 在零样本条件下实现3D姿态与语义对应估计,性能显著优于现有方法。
  • 适合对自监督3D建模、通用物体理解感兴趣的开发者和研究者。

3D可变形模型(3DMM)是表示特定物体类别可能形状与外观的强大工具。给定单张测试图像,3DMM可用于预测物体的3D形状、姿态、语义对应和实例分割等任务。然而,目前仅少数感兴趣类别(如人脸或人体)拥有可用的3DMM,因其需大量3D数据采集及类别专属训练。为此,我们提出Common3D,一种完全自监督的方法,从物体中心视频集合中学习常见物体的3DMM。模型将物体表示为学习得到的3D模板网格与由图像条件神经网络参数化的形变场。不同于以往工作,Common3D使用神经特征而非RGB颜色表示物体外观,通过像素强度抽象获得更具泛化性的表征。关键在于,利用可变形模板定义的对应关系,采用对比损失训练外观特征,从而获得更高质量的对应特征,显著提升3D姿态与语义对应估计性能。Common3D是首个能以零样本方式解决多种视觉任务的完全自监督方法。

原文摘要 · Abstract (English)

3D morphable models (3DMMs) are a powerful tool to represent the possible shapes and appearances of an object category. Given a single test image, 3DMMs can be used to solve various tasks, such as predicting the 3D shape, pose, semantic correspondence, and instance segmentation of an object. Unfortunately, 3DMMs are only available for very few object categories that are of particular interest, like faces or human bodies, as they require a demanding 3D data acquisition and category-specific training process. In contrast, we introduce a new method, Common3D, that learns 3DMMs of common objects in a fully self-supervised manner from a collection of object-centric videos. For this purpose, our model represents objects as a learned 3D template mesh and a deformation field that is parameterized as an image-conditioned neural network. Different from prior works, Common3D represents the object appearance with neural features instead of RGB colors, which enables the learning of more generalizable representations through an abstraction from pixel intensities. Importantly, we train the appearance features using a contrastive objective by exploiting the correspondences defined through the deformable template mesh. This leads to higher quality correspondence features compared to related works and a significantly improved model performance at estimating 3D object pose and semantic correspondence. Common3D is the first completely self-supervised method that can solve various vision tasks in a zero-shot manner.

3D建模自监督学习神经特征零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。