统一理解物体3D朝向与旋转,支持复杂对称性与相对旋转预测。
Orient Anything V2: Unifying Orientation and Rotation Understanding
- 用生成模型合成海量3D资产,覆盖广泛类别并均衡数据分布。
- 可识别0到N个有效正面,直接输出相对旋转,准确率领先基准。
- 适合需要精准朝向理解的机器人、三维重建等场景。
本文提出Orient Anything V2,一个增强型基础模型,可从单图或配对图像中统一理解物体的3D朝向与旋转。相较于V1仅定义单一唯一正面,V2扩展至处理具有多样化旋转对称性的物体,并直接估计相对旋转。该提升基于四项关键创新:1)利用生成模型合成大规模3D资产,确保类别覆盖广且数据分布均衡;2)设计高效“模型内标注”系统,可鲁棒识别每个物体的0至N个有效正面;3)提出对称感知的周期分布拟合目标,捕捉所有可能的正面朝向,有效建模物体旋转对称性;4)采用多帧架构,直接预测物体间的相对旋转。大量实验表明,该模型在11个主流基准上实现零样本朝向估计、6DoF姿态估计及物体对称性识别的最先进性能,展现出强泛化能力,显著拓展了朝向估计在各类下游任务中的应用范围。
原文摘要 · Abstract (English)
This work presents Orient Anything V2, an enhanced foundation model for unified understanding of object 3D orientation and rotation from single or paired images. Building upon Orient Anything V1, which defines orientation via a single unique front face, V2 extends this capability to handle objects with diverse rotational symmetries and directly estimate relative rotations. These improvements are enabled by four key innovations: 1) Scalable 3D assets synthesized by generative models, ensuring broad category coverage and balanced data distribution; 2) An efficient, model-in-the-loop annotation system that robustly identifies 0 to N valid front faces for each object; 3) A symmetry-aware, periodic distribution fitting objective that captures all plausible front-facing orientations, effectively modeling object rotational symmetry; 4) A multi-frame architecture that directly predicts relative object rotations. Extensive experiments show that Orient Anything V2 achieves state-of-the-art zero-shot performance on orientation estimation, 6DoF pose estimation, and object symmetry recognition across 11 widely used benchmarks. The model demonstrates strong generalization, significantly broadening the applicability of orientation estimation in diverse downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。