首个农作物6D位姿与形变联合标注基准,提升机器人采摘精度。
Mind the Shape Gap: A Benchmark and Baseline for Deformation-Aware 6D Pose Estimation of Agricultural Produce
- 提出统一框架SEED,从单张图像同时估计位姿与显式形变。
- 在8类作物上实测现役方法性能下降达6倍,凸显形变建模重要性。
- 仅用合成数据训练,适用于多品类农业场景的高鲁棒性位姿估计。
机器人采摘中的精确6D位姿估计因农产品的生物可变形性及类内形状高度变异而受阻。实例级方法难以应用,因难以获取每件农产品的精确3D模型;类别级方法依赖固定模板,在先验与真实几何偏差时性能显著下降。为弥补这一缺陷,我们引入PEAR(Pose and dEformation of Agricultural pRoduce),首个跨8类农产品提供联合6D位姿与实例级3D形变真值的基准,通过机械臂实现高精度标注。利用PEAR,我们发现现有最先进方法在真实农产品几何偏差下性能最高下降6倍。基于此,我们提出SEED(Simultaneous Estimation of posE and Deformation),一种统一的仅含RGB输入的框架,可跨多个作物类别从单张图像联合预测6D位姿与显式网格形变。该模型完全在合成数据上训练,并在UV层面应用生成纹理增强,其在8类中的6类上超越MegaPose,证明显式形状建模是实现农业机器人可靠位姿估计的关键。
原文摘要 · Abstract (English)
Accurate 6D pose estimation for robotic harvesting is fundamentally hindered by the biological deformability and high intra-class shape variability of agricultural produce. Instance-level methods fail in this setting, as obtaining exact 3D models for every unique piece of produce is practically infeasible, while category-level approaches that rely on a fixed template suffer significant accuracy degradation when the prior deviates from the true instance geometry. To bridge such lack of robustness to deformation, we introduce PEAR (Pose and dEformation of Agricultural pRoduce), the first benchmark providing joint 6D pose and per-instance 3D deformation ground truth across 8 produce categories, acquired via a robotic manipulator for high annotation accuracy. Using PEAR, we show that state-of-the-art methods suffer up to 6x performance degradation when faced with the inherent geometric deviations of real-world produce. Motivated by this finding, we propose SEED (Simultaneous Estimation of posE and Deformation), a unified RGB-only framework that jointly predicts 6D pose and explicit lattice deformations from a single image across multiple produce categories. Trained entirely on synthetic data with generative texture augmentation applied at the UV level, SEED outperforms MegaPose on 6 out of 8 categories under identical RGB-only conditions, demonstrating that explicit shape modeling is a critical step toward reliable pose estimation in agricultural robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。