arXiv:2605.00307cs.ROcs.CV2026-05被引 1

用视觉+物理模型估算软夹爪受力,实时准确且适应新物体。

A Model-based Visual Contact Localization and Force Sensing System for Compliant Robotic Grippers

论文配图:A Model-based Visual Contact Localization and Force Sensing System for Compliant Robotic Grippers
图 1 · 摘自论文原文
  • 通过相机捕捉变形夹爪的结构点,驱动有限元仿真反推受力。
  • 负载阶段误差仅0.23N,全程误差0.48N,精度高且稳定。
  • 适合需要精准力控的柔性抓取场景,如抓握易损物品。

抓握力估计有助于防止机器人在操作中损坏脆弱物体,并提升基于学习的机器人控制性能。将力传感集成到可变形夹爪中需权衡成本、复杂度、机械鲁棒性与性能。随着RGB-D腕部相机在机器人系统中日益普及,基于视觉的间接力估计成为有前景的解决方案。现有方法多采用端到端深度学习,泛化能力弱;而传统模型基方法不适用于现代夹爪结构和抓取任务。为此,我们提出一种融合迭代接触定位的模型基视觉力感知方法,可推广至未见物体。系统从腕部相机获取的RGB-D图像中提取软夹爪的结构关键点,用于定义仿真开放框架架构(SOFA)中的逆有限元分析参数。迭代接触定位子系统采用基于深度学习的在线三维重建与位姿估计算法,动态更新接触位置,对视觉遮挡和未知物体具有鲁棒性。实验表明,在不同物体与条件下,系统在加载阶段平均均方根误差为0.23 N,归一化均方根偏差为2.11%;整个抓取过程分别为0.48 N和4.34%,展现出实时模型基间接力感知的潜力。

原文摘要 · Abstract (English)

Grasp force estimation can help prevent robots from damaging delicate objects during manipulation and improve learning-based robotic control. Integrating force sensing into deformable grippers negotiates trade-offs in cost, complexity, mechanical robustness, and performance. With the growing integration of RGB-D wrist cameras into robotic systems for control purposes, camera-based techniques are a promising solution for indirect visual force estimation. Current approaches mostly utilize end-to-end deep learning, which can be brittle when generalizing to new scenarios, while existing model-based approaches are unsuited to grasping and modern grasper geometries. To address these challenges, we developed a model-based visual force sensing approach integrating an iterative contact localization with generalization to unseen objects. The system extracts structural key points from wrist camera RGB-D images of deforming fin-ray-shaped soft grippers, and uses these key points to define parameters of an inverse finite element analysis simulation in Simulation Open Framework Architecture. The iterative contact localization sub-system utilizes a deep learning-based online 3D reconstruction and pose estimation pipeline to dynamically update contact location, and is robust to visual occlusion and unseen objects. Our system demonstrated an average root mean square error of 0.23 N and normalized root mean square deviation of 2.11% during the load phase, and 0.48 N and 4.34% over the entire grasping process when interacting with different objects under various conditions, showcasing its potential for real-time model-based indirect force sensing of soft grippers.

力感知软体机器人视觉估计有限元仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。