arXiv:2411.00543cs.CVcs.LG2024-11NeurIPS被引 6

直接预测旋转系数,实现更精准的3D姿态估计。

3D Equivariant Pose Regression via Direct Wigner-D Harmonics Prediction

  • 在频域直接预测威格纳-D系数,避免空间参数化缺陷。
  • 在ModelNet10-SO(3)和PASCAL3D+上达到领先性能。
  • 适合需要高精度、抗旋转变化的3D视觉任务。

单图像3D姿态估计是3D视觉中的关键任务。现有方法多采用欧拉角或四元数等空间域参数化表示3D旋转,但易引入不连续性和奇点。SO(3)-等变网络能高效捕获姿态模式,但其架构(如球面卷积神经网络)运行于频域,与空间参数化不兼容。为此,本文提出一种频域方法,直接预测3D旋转的威格纳-D系数,与球面CNN操作对齐。所提出的SO(3)-等变姿态谐波预测器克服了空间参数化的局限,确保任意旋转下的稳定估计。通过频域回归损失训练,在ModelNet10-SO(3)和PASCAL3D+等基准上取得当前最优结果,显著提升准确率、鲁棒性与数据效率。

原文摘要 · Abstract (English)

Determining the 3D orientations of an object in an image, known as single-image pose estimation, is a crucial task in 3D vision applications. Existing methods typically learn 3D rotations parametrized in the spatial domain using Euler angles or quaternions, but these representations often introduce discontinuities and singularities. SO(3)-equivariant networks enable the structured capture of pose patterns with data-efficient learning, but the parametrizations in spatial domain are incompatible with their architecture, particularly spherical CNNs, which operate in the frequency domain to enhance computational efficiency. To overcome these issues, we propose a frequency-domain approach that directly predicts Wigner-D coefficients for 3D rotation regression, aligning with the operations of spherical CNNs. Our SO(3)-equivariant pose harmonics predictor overcomes the limitations of spatial parameterizations, ensuring consistent pose estimation under arbitrary rotations. Trained with a frequency-domain regression loss, our method achieves state-of-the-art results on benchmarks such as ModelNet10-SO(3) and PASCAL3D+, with significant improvements in accuracy, robustness, and data efficiency.

3D姿态估计等变网络频域建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。