arXiv:2608.00464cs.RO2026-08

让机器人从视觉预测物体接触力分布,提升抓取稳定性。

Learning to Predict Contact Force Distributions from Vision Leveraging Object Geometry Priors

论文配图:Learning to Predict Contact Force Distributions from Vision Leveraging Object Geometry Priors
图 1 · 摘自论文原文
  • 用单张图像预测三维力分布,而非原始点力。
  • 结合物体几何信息平滑力分布,精度显著提升。
  • 模型仅在仿真训练,却能有效泛化到真实场景。

人类基于视觉和经验可进行粗略的物理预测并调整操作策略。本文旨在赋予机器人类似能力。为获取视觉与力的配对数据,采用机器人领域常用的刚体模拟器。不同于输出噪声点力的模拟器,人类即使在陌生情境下也能做出一致预测。基于此,我们假设预测平滑的力分布而非原始点力,可同时提升预测准确率与下游任务表现。为此,构建了一个从单张RGB图像预测三维力分布的模型。目标分布通过将模拟器生成的点力进行统计平滑得到。进一步地,将物体几何信息融入平滑过程,以适应不同接触状态,实现更一致的视觉预测。在仿真和真实环境中的大量实验表明,该方法提升了预测精度,平滑机制增强了下游任务性能,几何引导的平滑效果更优。令人惊讶的是,模型虽仅在仿真中训练,却能有效泛化至真实场景。

原文摘要 · Abstract (English)

Based on vision and prior experience, humans can make rough physical predictions and adjust their manipulation strategies. This paper aims to endow robots with a similar ability. To collect paired data of vision and forces, we use a rigid-body simulator commonly adopted in robotics. However, unlike simulators that output noisy point forces, humans are able to make consistent predictions even in unfamiliar situations. Based on this observation, we hypothesize that predicting smooth force distributions rather than raw point forces can improve both force prediction itself and downstream task performance. To validate this hypothesis, we construct a model that predicts three-dimensional force distributions from a single RGB image of piled daily objects. The target distribution is generated by applying statistical smoothing to point forces obtained from the simulator. Moreover, by incorporating object geometry into the smoothing process, we aim to account for variations in contact states and achieve more consistent vision-based predictions. We conduct extensive evaluations in both simulation and real environments. Results show that our approach improves prediction accuracy, enhances downstream task performance through smoothing, and further benefits from geometry-guided smoothing. Remarkably, the trained model generalizes effectively to real-world scenes despite being trained solely in simulation.

力预测视觉感知机器人操作仿真泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。