用自然语言指令控制机械手抓取,通用性强且动作真实
UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
- 将不同形态机械手统一编码,提升跨模型泛化能力
- 仅用人类交互数据训练,生成动作自然且符合物理规律
- 适合需要多任务、开放指令的机器人操控场景
规划具备灵巧性的机械手操作是机器人操纵与具身智能的核心挑战。以往方法依赖物体中心线索或精确的手物交互序列,忽视了开放词汇指令中丰富的组合性指导。我们提出UniHM,首个基于自由格式语言指令的统一灵巧手操作框架。设计统一手部灵巧编码器,将异构灵巧手形态映射到单一共享代码本,显著提升跨形态泛化与可扩展性。视觉语言动作模型仅在人类-物体交互数据上训练,无需大量真实遥操作数据,即可从开放式语言指令生成类人操作序列。为确保物理真实性,引入物理引导的动态优化模块,在生成与时间先验下进行分段关节优化,获得平滑且物理可行的操作序列。在多个数据集和真实世界评估中,UniHM在已见与未见物体及轨迹上均达到最优表现,展现出强大泛化能力与高物理可行性。
原文摘要 · Abstract (English)
Planning physically feasible dexterous hand manipulation is a central challenge in robotic manipulation and Embodied AI. Prior work typically relies on object-centric cues or precise hand-object interaction sequences, foregoing the rich, compositional guidance of open-vocabulary instruction. We introduce UniHM, the first framework for unified dexterous hand manipulation guided by free-form language commands. We propose a Unified Hand-Dexterous Tokenizer that maps heterogeneous dexterous-hand morphologies into a single shared codebook, improving cross-dexterous hand generalization and scalability to new morphologies. Our vision language action model is trained solely on human-object interaction data, eliminating the need for massive real-world teleoperation datasets, and demonstrates strong generalizability in producing human-like manipulation sequences from open-ended language instructions. To ensure physical realism, we introduce a physics-guided dynamic refinement module that performs segment-wise joint optimization under generative and temporal priors, yielding smooth and physically feasible manipulation sequences. Across multiple datasets and real-world evaluations, UniHM attains state-of-the-art results on both seen and unseen objects and trajectories, demonstrating strong generalization and high physical feasibility. Our project page at \href{https://unihm.github.io/}{https://unihm.github.io/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。