arXiv:2510.17086cs.RO2025-10被引 4

用奖励模型优化软体手设计,提升抓取成功率

Learning to Design Soft Hands using Reward Models

  • 基于奖赏模型的交叉熵方法高效搜索最优软体手结构
  • 实测抓取成功率显著高于基线,且减少一半以上设计评估次数
  • 适合需要快速迭代软体机器人设计的研究者与工程师

软体机器人手有望实现与物体和环境的安全柔顺交互。然而,设计出在多种场景下兼具柔顺性与功能性的软体手仍具挑战。尽管硬件与控制协同设计能更好匹配形态与行为,但其搜索空间维度高,即使基于仿真的评估也极为耗时。本文提出一种基于奖赏模型的交叉熵方法(CEM-RM),通过遥操作控制策略优化肌腱驱动的软体手,相比纯优化方法减少超过一半的设计评估次数,并从预先收集的遥操作数据中学习到一组优化的手部设计分布。我们构建了由弯曲软指组成的软体手设计空间,并在仿真中实现并行训练。优化后的手部通过3D打印实现,分别使用遥操作数据和实时遥操作在真实世界部署。仿真与硬件实验均表明,优化设计在多样化挑战性物体上的抓取成功率显著优于基线手。

原文摘要 · Abstract (English)

Soft robotic hands promise to provide compliant and safe interaction with objects and environments. However, designing soft hands to be both compliant and functional across diverse use cases remains challenging. Although co-design of hardware and control better couples morphology to behavior, the resulting search space is high-dimensional, and even simulation-based evaluation is computationally expensive. In this paper, we propose a Cross-Entropy Method with Reward Model (CEM-RM) framework that efficiently optimizes tendon-driven soft robotic hands based on teleoperation control policy, reducing design evaluations by more than half compared to pure optimization while learning a distribution of optimized hand designs from pre-collected teleoperation data. We derive a design space for a soft robotic hand composed of flexural soft fingers and implement parallelized training in simulation. The optimized hands are then 3D-printed and deployed in the real world using both teleoperation data and real-time teleoperation. Experiments in both simulation and hardware demonstrate that our optimized design significantly outperforms baseline hands in grasping success rates across a diverse set of challenging objects.

软体机器人强化学习设计优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。