让机械手在复杂交互中更智能地操作,避免虚幻动作,提升成功率。
DexHandDiff: Interaction-aware Diffusion Planning for Adaptive Dexterous Manipulation
- 用双阶段扩散模型模拟接触前对齐与接触后控制,增强适应性。
- 在多个任务上成功率超60%,比现有方法高一倍以上。
- 支持语言指令引导,适合需要灵活操作的机器人研究者。
具有丰富接触交互的灵巧操作对先进机器人至关重要。尽管基于扩散的规划方法在简单任务中表现良好,但在处理复杂序列交互时常出现不现实的‘幽灵状态’(如物体自动移动而无手接触),且缺乏适应性。本文提出DexHandDiff,一种面向自适应灵巧操作的交互感知扩散规划框架。该框架通过双阶段扩散过程建模联合状态-动作动态,包括接触前对齐与接触后目标导向控制,实现目标自适应、可泛化的灵巧操作。此外,引入基于动力学模型的双重引导,并利用大语言模型自动生成引导函数,增强物理交互的泛化能力,通过语言提示实现多样化目标适配。在门开启、笔与方块翻转、物体重定位及锤击等物理交互任务上的实验表明,DexHandDiff在训练分布外的目标上表现优异,平均成功率(59.2%)超过现有方法(29.5%)一倍以上;在目标自适应灵巧操作任务中平均成功率达70.7%,凸显其在高接触密度操作中的鲁棒性与灵活性。
原文摘要 · Abstract (English)
Dexterous manipulation with contact-rich interactions is crucial for advanced robotics. While recent diffusion-based planning approaches show promise for simple manipulation tasks, they often produce unrealistic ghost states (e.g., the object automatically moves without hand contact) or lack adaptability when handling complex sequential interactions. In this work, we introduce DexHandDiff, an interaction-aware diffusion planning framework for adaptive dexterous manipulation. DexHandDiff models joint state-action dynamics through a dual-phase diffusion process which consists of pre-interaction contact alignment and post-contact goal-directed control, enabling goal-adaptive generalizable dexterous manipulation. Additionally, we incorporate dynamics model-based dual guidance and leverage large language models for automated guidance function generation, enhancing generalizability for physical interactions and facilitating diverse goal adaptation through language cues. Experiments on physical interaction tasks such as door opening, pen and block re-orientation, object relocation, and hammer striking demonstrate DexHandDiff's effectiveness on goals outside training distributions, achieving over twice the average success rate (59.2% vs. 29.5%) compared to existing methods. Our framework achieves an average of 70.7% success rate on goal adaptive dexterous tasks, highlighting its robustness and flexibility in contact-rich manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。