arXiv:2608.30773cs.RO2026-08

让软体机器人通过全身接触感知和推理,实现盲抓取

Learning to infer and manipulate through distributed whole-arm interaction in a soft robot

论文配图:Learning to infer and manipulate through distributed whole-arm interaction in a soft robot
图 1 · 摘自论文原文
  • 用分布式柔性接触同时获取信息并组织抓取行为
  • 在无视觉条件下成功抓取多种物体,依赖接触历史推理
  • 适合对物理智能、软体机器人感兴趣的学者

在象类和章鱼等动物中,获取物体的非视觉信息与物理交互密不可分,由柔韧附肢与环境的大面积互动所驱动。软体机器人为将这一原理转化为工程系统提供了天然平台。然而当前机器人智能对物理交互利用有限,通常将其视为需规避的干扰或仅用于补偿物体错位。本文提出一种物理智能框架,使分布式的柔顺交互共同揭示任务相关的信息并组织操作行为。这导致一个固有的部分可观测问题:关键任务信息从未被直接测量,而必须从物理交互的历史中推断。我们设计了一种基于强化学习的架构,通过端到端学习记忆型控制策略来应对这一挑战。核心创新包括:(i) 预训练探索策略提供广域工作空间探索参考;(ii) 联合优化将探索与抓取目标整合进单一循环策略;(iii) 两阶段仿真到现实迁移,包含观测映射与策略微调。我们通过配备嵌入式惯性测量单元(IMUs)于柔顺结构中的混合刚-软机械臂,验证了该原理。该学习策略能自主协调工作空间探索、物体相遇与定位、抓取相关属性推断及稳定全身包裹,实现盲抓取。

原文摘要 · Abstract (English)

In animals such as elephants and octopuses, acquiring non-visual information about an object and physically engaging with it are inseparable processes mediated by rich, large-area interactions between compliant appendages and the environment. Soft robots provide a natural platform for translating this principle into engineered systems. Yet current robotic intelligence makes limited use of physical interaction, treating it primarily as a disturbance to be rejected or, at best, as a means of compensating for object misalignment. Here, we introduce a physical intelligence framework in which distributed compliant interactions jointly reveal task-relevant information and organize manipulation behavior. This results in an intrinsically partially observable problem: key task-relevant information is never measured directly, but must instead be inferred from the history of physical interactions. We propose a reinforcement-learning architecture that addresses this challenge by learning a memory-based control policy end-to-end. The key innovations making this possible are (i) a pretrained exploration policy that provides a reference for broad workspace exploration, (ii) joint optimization that integrates exploration and grasping objectives within a single recurrent policy, and (iii) a two-stage sim-to-real adaptation including observation mapping and policy fine-tuning. We demonstrate this principle through blind whole-arm grasping with a hybrid rigid-soft robotic arm that we equip with IMUs embedded directly within its compliant structure, providing its only source of proprioceptive sensing. The learned policy successfully identifies and grasps various objects by autonomously coordinating workspace exploration, object encounter and localization, inference of grasp-relevant properties, and stable whole-arm wrapping.

软体机器人物理智能强化学习盲抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。