用视觉语言模型从过往经验中预测抓取初始力,避免压坏或滑脱。
Exp-Force: Experience-Conditioned Pre-Grasp Force Selection with Vision-Language Models
- 基于视觉-语言模型,结合少量历史抓取经验进行上下文推理。
- 在129个物体上误差降低72%,真实场景中正确率从63%提升至87%。
- 无需物理建模,适用于柔性夹爪等难建模场景,适合机器人抓取任务。
准确的接触前抓取力选择对安全可靠的机器人操作至关重要。自适应控制器虽可在接触后调节力,但仍需合理的初始估计。初始力过小需后续调整,过大则可能损坏脆弱物体。这一权衡对难以解析建模的柔性夹爪尤为困难。本文提出Exp-Force,一种基于经验的框架,仅凭单张RGB图像预测最小可行抓取力。该方法检索少量相关历史抓取经验,并将视觉-语言模型基于这些实例进行上下文推理,无需解析接触模型或人工设计启发式规则。在129个物体实例上,Exp-Force最佳情况下的平均绝对误差(MAE)为0.43 N,相比零样本推理减少72%误差。在30个未见过的真实物体测试中,合理力选择成功率从63%提升至87%。结果表明,Exp-Force通过利用过往交互经验,实现了可靠且可泛化的接触前抓取力选择。
原文摘要 · Abstract (English)
Accurate pre-contact grasp force selection is critical for safe and reliable robotic manipulation. Adaptive controllers regulate force after contact but still require a reasonable initial estimate. Starting a grasp with too little force requires reactive adjustment, while starting a grasp with too high a force risks damaging fragile objects. This trade-off is particularly challenging for compliant grippers, whose contact mechanics are difficult to model analytically. We propose Exp-Force, an experience-conditioned framework that predicts the minimum feasible grasping force from a single RGB image. The method retrieves a small set of relevant prior grasping experiences and conditions a vision-language model on these examples for in-context inference, without analytic contact models or manually designed heuristics. On 129 object instances, ExpForce achieves a best-case MAE of 0.43 N, reducing error by 72% over zero-shot inference. In real-world tests on 30 unseen objects, it improves appropriate force selection rate from 63% to 87%. These results demonstrate that Exp-Force enables reliable and generalizable pre-grasp force selection by leveraging prior interaction experiences. http://expforcesubmission.github.io/Exp-Force-Website/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。