用大模型设计触觉奖励,让机械手更灵活地翻转物体。
Text2Touch: Tactile In-Hand Manipulation with LLM-Designed Reward Functions
- 用大模型生成触觉反馈的奖励函数,自动优化抓握动作。
- 在真实触觉传感器下实现多轴旋转,速度和稳定性优于人工设计。
- 只需简短指令即可生成有效策略,适合快速部署复杂操作。
大型语言模型(LLMs)正开始自动化灵巧操作中的奖励设计,但以往工作未考虑触觉感知——这对类人灵巧性至关重要。我们提出Text2Touch,将大模型设计的奖励引入基于视觉触觉传感的真实世界多轴掌上物体旋转任务,覆盖掌心朝上与掌心朝下两种姿态。通过提示工程扩展至70多个环境变量,结合仿真到现实的迁移策略,成功将策略部署于配备触觉传感器的四指灵巧机械手。Text2Touch显著优于精心调优的人工奖励基线,旋转速度更快、更稳定,且所用奖励函数长度和复杂度低一个数量级。结果表明,大模型设计的奖励能极大缩短从概念到可部署灵巧触觉技能的时间,推动多模态机器人学习的快速与可扩展发展。
原文摘要 · Abstract (English)
Large language models (LLMs) are beginning to automate reward design for dexterous manipulation. However, no prior work has considered tactile sensing, which is known to be critical for human-like dexterity. We present Text2Touch, bringing LLM-crafted rewards to the challenging task of multi-axis in-hand object rotation with real-world vision based tactile sensing in palm-up and palm-down configurations. Our prompt engineering strategy scales to over 70 environment variables, and sim-to-real distillation enables successful policy transfer to a tactile-enabled fully actuated four-fingered dexterous robot hand. Text2Touch significantly outperforms a carefully tuned human-engineered baseline, demonstrating superior rotation speed and stability while relying on reward functions that are an order of magnitude shorter and simpler. These results illustrate how LLM-designed rewards can significantly reduce the time from concept to deployable dexterous tactile skills, supporting more rapid and scalable multimodal robot learning. Project website: https://hpfield.github.io/text2touch-website
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。