用3D渲染数据训练模型,实现单图物体朝向精准估计
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
- 从3D模型渲染200万张带朝向标注的图像,构建大规模训练集
- 将朝向建模为三维角度的概率分布,提升估计鲁棒性
- 零样本适配多种场景,适用于3D姿态调整与空间理解任务
朝向是物体的关键属性,对理解其在图像中的空间姿态和排列至关重要。然而,从单张图像中准确估计朝向的实际解决方案仍不充分。本文提出Orient Anything,首个专为单视角和任意视角图像设计的物体朝向估计基础模型。由于标注数据稀缺,我们通过从3D世界提取知识来解决:构建管线自动标注3D物体的正面,并从随机视角渲染图像,生成包含200万张带精确朝向标注的图像数据集。为充分利用该数据集,我们设计了一种鲁棒的训练目标,将3D朝向建模为三个角度的概率分布,并通过拟合这些分布来预测物体朝向。此外,采用多种策略提升合成数据到真实世界的迁移性能。模型在渲染图像和真实图像上均达到当前最佳朝向估计精度,并展现出出色的零样本泛化能力。更重要的是,该模型显著提升了复杂空间概念理解与生成、3D物体姿态调整等应用效果。
原文摘要 · Abstract (English)
Orientation is a key attribute of objects, crucial for understanding their spatial pose and arrangement in images. However, practical solutions for accurate orientation estimation from a single image remain underexplored. In this work, we introduce Orient Anything, the first expert and foundational model designed to estimate object orientation in a single- and free-view image. Due to the scarcity of labeled data, we propose extracting knowledge from the 3D world. By developing a pipeline to annotate the front face of 3D objects and render images from random views, we collect 2M images with precise orientation annotations. To fully leverage the dataset, we design a robust training objective that models the 3D orientation as probability distributions of three angles and predicts the object orientation by fitting these distributions. Besides, we employ several strategies to improve synthetic-to-real transfer. Our model achieves state-of-the-art orientation estimation accuracy in both rendered and real images and exhibits impressive zero-shot ability in various scenarios. More importantly, our model enhances many applications, such as comprehension and generation of complex spatial concepts and 3D object pose adjustment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。