arXiv:2503.06773cs.CV2025-03

研究3D物体图像的低维流形结构,揭示其几何特性与分类关联。

Investigating Image Manifolds of 3D Objects: Learning, Shape Analysis, and Comparisons

  • 用保几何变换将物体姿态图像映射到低维空间
  • 发现不同物体的图像流形为光滑非线性结构
  • 基于形状分析实现物体分类聚类,适用于视觉模型研究

尽管图像维度高,但3D物体的图像集合长期被认为构成低维流形。本文从新几何视角重新审视流形学习问题,利用保几何变换将物体姿态图像流形映射至低维潜在空间。结果显示,不同物体在潜在空间中的姿态流形均为光滑、非线性的低维流形。进一步采用Kendall形状分析(模去刚体运动与全局缩放)比较各物体流形形状,并据此聚类。有趣的是,同类别物体的流形常被聚在一起。这些图像流形的几何结构可被用于简化视觉任务、预测性能,并深化对学习方法的理解。

原文摘要 · Abstract (English)

Despite high-dimensionality of images, the sets of images of 3D objects have long been hypothesized to form low-dimensional manifolds. What is the nature of such manifolds? How do they differ across objects and object classes? Answering these questions can provide key insights in explaining and advancing success of machine learning algorithms in computer vision. This paper investigates dual tasks -- learning and analyzing shapes of image manifolds -- by revisiting a classical problem of manifold learning but from a novel geometrical perspective. It uses geometry-preserving transformations to map the pose image manifolds, sets of images formed by rotating 3D objects, to low-dimensional latent spaces. The pose manifolds of different objects in latent spaces are found to be nonlinear, smooth manifolds. The paper then compares shapes of these manifolds for different objects using Kendall's shape analysis, modulo rigid motions and global scaling, and clusters objects according to these shape metrics. Interestingly, pose manifolds for objects from the same classes are frequently clustered together. The geometries of image manifolds can be exploited to simplify vision and image processing tasks, to predict performances, and to provide insights into learning methods.

流形学习3D视觉几何分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。