arXiv:2409.16600cs.CV2024-09ECCV被引 3

无需真实标注,仅用3D模型和图像实现水下物体6自由度姿态估计。

FAFA: Frequency-Aware Flow-Aided Self-Supervision for Underwater Object Pose Estimation

  • 基于频域感知的光流辅助自监督学习,提升水下环境适应性。
  • 在常见水下基准上显著优于现有方法,性能大幅提升。
  • 适合水下机器人、海洋探测等缺乏标注数据的场景使用。

尽管室内场景中物体姿态估计已取得显著进展,但复杂水下环境带来的光照衰减、模糊及真实标注成本高昂等问题,使得水下物体姿态估计仍具挑战。为此,我们提出FAFA——一种面向无人水下航行器(UUV)6D姿态估计的频域感知光流辅助自监督框架。首先,在合成数据上训练一个频域感知的光流姿态估计器,提出基于FFT的增强方法,帮助网络从频域角度捕捉域不变特征与目标域风格。随后,通过强制执行光流辅助的多层级一致性进行自监督训练,使模型适配真实水下环境。本方法仅依赖3D模型与RGB图像,无需任何真实姿态标注或其他模态数据(如深度)。我们在常用水下物体姿态估计基准上评估了FAFA的有效性,结果表明其显著优于当前最先进方法。代码已开源。

原文摘要 · Abstract (English)

Although methods for estimating the pose of objects in indoor scenes have achieved great success, the pose estimation of underwater objects remains challenging due to difficulties brought by the complex underwater environment, such as degraded illumination, blurring, and the substantial cost of obtaining real annotations. In response, we introduce FAFA, a Frequency-Aware Flow-Aided self-supervised framework for 6D pose estimation of unmanned underwater vehicles (UUVs). Essentially, we first train a frequency-aware flow-based pose estimator on synthetic data, where an FFT-based augmentation approach is proposed to facilitate the network in capturing domain-invariant features and target domain styles from a frequency perspective. Further, we perform self-supervised training by enforcing flow-aided multi-level consistencies to adapt it to the real-world underwater environment. Our framework relies solely on the 3D model and RGB images, alleviating the need for any real pose annotations or other-modality data like depths. We evaluate the effectiveness of FAFA on common underwater object pose benchmarks and showcase significant performance improvements compared to state-of-the-art methods. Code is available at github.com/tjy0703/FAFA.

姿态估计自监督学习水下视觉频域增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。