arXiv:2505.24636cs.CV2025-05被引 1

用RGB图像实现农业果蔬的6维姿态估计,无需特定模型

Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic Data

  • 基于基底网格变形预测姿态与形变参数
  • 合成数据经扩散模型增强纹理,提升真实感
  • 在香蕉多形态数据集上显著优于现有方法

精确的6D物体姿态估计对农业机器人抓取至关重要,但果蔬在形状、大小和纹理上存在高度类内差异。现有方法大多依赖实例级CAD模型或深度传感器,难以用于真实农业场景。本文提出PLANTPose框架,仅使用RGB输入即可实现类别级6D姿态估计。该框架预测相对于基底网格的6D姿态与形变参数,使单一类别级CAD模型可适应未见实例,无需实例特定数据。为提升合成数据的真实感并增强泛化能力,我们采用Stable Diffusion对合成图像进行纹理细化,模拟成熟度与环境因素带来的变化,缩小合成数据与真实世界的域差距。在包含多种形状、大小和成熟度香蕉的挑战性基准测试中,该框架有效应对大类内变异,保持高精度6D姿态预测,显著优于当前最先进的基于RGB的方法MegaPose。

原文摘要 · Abstract (English)

Accurate 6D object pose estimation is essential for robotic grasping and manipulation, particularly in agriculture, where fruits and vegetables exhibit high intra-class variability in shape, size, and texture. The vast majority of existing methods rely on instance-specific CAD models or require depth sensors to resolve geometric ambiguities, making them impractical for real-world agricultural applications. In this work, we introduce PLANTPose, a novel framework for category-level 6D pose estimation that operates purely on RGB input. PLANTPose predicts both the 6D pose and deformation parameters relative to a base mesh, allowing a single category-level CAD model to adapt to unseen instances. This enables accurate pose estimation across varying shapes without relying on instance-specific data. To enhance realism and improve generalization, we also leverage Stable Diffusion to refine synthetic training images with realistic texturing, mimicking variations due to ripeness and environmental factors and bridging the domain gap between synthetic data and the real world. Our evaluations on a challenging benchmark that includes bananas of various shapes, sizes, and ripeness status demonstrate the effectiveness of our framework in handling large intraclass variations while maintaining accurate 6D pose predictions, significantly outperforming the state-of-the-art RGB-based approach MegaPose.

6D姿态估计农业机器人扩散模型类别级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。