arXiv:2607.07897cs.RO2026-07中稿 · IEEE/ASME Internat…

仅用单目视觉和位置控制夹爪,统一抓取软硬物体。

Monocular Vision Based Control Framework for Grasping

论文配图:Monocular Vision Based Control Framework for Grasping
图 1 · 摘自论文原文
  • 基于语义描述估计物体刚度,决定抓取策略。
  • 通过关键点追踪与变形度量实现软物自适应抓取。
  • 实测在5类差异大的物体上稳定抓取,无需触觉传感器。

在非结构化环境中抓取物体需应对从柔软易变形到坚硬日常物品的广泛机械特性差异。现有方法通常分治处理,依赖触觉感知、对象特定模型或专用夹爪。本文提出一种统一的单目视觉抓取框架,仅使用RGB输入与位置控制夹爪,即可同时处理软硬物体。系统融合开放词汇目标检测、图像分割、边界感知点分配、实时点跟踪及单目深度估计,从视觉观测中恢复物体运动与几何信息。核心组件为语言驱动的刚度估计模型,根据语义描述推断物体预期柔顺性,并在接触前提供抓取策略先验。对于可变形物体,通过追踪关键点计算的Procrustes相似性度量作为形变视觉代理,指导抓取适应;对于刚性物体,通过追踪点间距离缩放调节夹爪宽度。在Franka Emika Research 3机械臂上对生菜、新鲜马苏里拉奶酪、牛角包、纸巾和硬塑料瓶等五类差异显著物体进行真实世界拾取-放置实验,结果表明该框架仅依靠视觉反馈即可在软硬物体间实现稳定抓取,展示了一种实用、传感器高效且通用的食品处理与家庭操作方案。

原文摘要 · Abstract (English)

Grasping in unstructured environments requires handling objects with widely different mechanical properties, from soft and deformable items to rigid everyday objects. Most existing approaches address these categories separately and often rely on tactile sensing, object-specific models, or specialized grippers. In this paper, we present a unified monocular vision-based grasping framework that targets both soft and rigid objects within a single control pipeline, using only RGB input and a position-controlled gripper. The proposed system combines open-vocabulary object detection, image segmentation, boundary-aware point assignment, real-time point tracking, and monocular depth estimation to recover object motion and geometry from visual observations. A key component of the framework is a language-based stiffness estimation model that infers an object's expected compliance from its semantic description and provides an object-level prior for selecting the grasping strategy before contact. For deformable objects, grasp adaptation is governed by a Procrustes-based dissimilarity measure computed from tracked keypoints, which acts as a visual proxy for deformation. For rigid objects, the gripper width is regulated through the scaling of tracked point distances. We validate the proposed method in real-world pick-and-place experiments on a Franka Emika Research 3 arm using objects with substantially different mechanical properties, including lettuce, fresh mozzarella cheese, croissants, paper towels, and hard plastic bottles. Results demonstrate that the framework achieves stable grasping across both soft and rigid objects using visual feedback alone, highlighting a practical, sensor-efficient, and generalizable approach for food handling and household manipulation.

单目视觉抓取柔性物体机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。