提出分层框架,让机器人用双手灵巧操作工具更高效准确。
DexFuture: Hierarchical Future-State Visuomotor Targeting for Bimanual Dexterous Tool Use

- 分层设计:高层预测未来动作目标,底层精细控制执行
- 在真实任务中达到90%理想表现,比无目标方法提升近13倍
- 每秒运行60次,比传统规划快250倍,适合实时应用
双臂灵巧操作工具对机器人仍具挑战,源于高维手部构型与复杂的手-工具-物体动态及接触关系。现有控制策略多依赖示范提供的未来配置参考,而基于未来动作的世界模型需进行耗时的在线规划。关键难点在于不依赖特权状态或缓慢反事实规划,生成动态一致的未来参考轨迹。本文提出DexFuture,一个分层系统,包含高层未来状态视觉运动目标预测器与底层目标条件化结构化灵巧策略。基于自车视角RGB、本体感知和几何历史,高层预测器构建结构化的手-工具-物体视觉运动嵌入,利用时序条件化的Transformer生成多步未来目标轨迹;底层策略则通过目标条件化的逐关节Transformer追踪这些目标。该分层架构解耦了粗粒度未来参考生成与细粒度动作控制,以及长时程语义预测与高频执行。在OakInk2双臂工具使用任务上,DexFuture达到特权-基准性能的90%,远超无参考策略的7%。系统运行频率达60 Hz,约为基于未来动作条件世界模型的CEM规划方法的250倍速度。
原文摘要 · Abstract (English)
Bimanual dexterous tool use remains challenging for robots due to high-dimensional hand configurations and complex hand-tool-object dynamics and contact. Most existing control policies depend on future configuration references provided from demonstrations, while future action-conditioned world models require slow online planning over high-dimensional action sequences. A significant challenge is generating a dynamically consistent future reference trajectory without relying on privileged states from demonstrations or slow counterfactual planning. We propose DexFuture, a hierarchical system that couples a high-level Future-State Visuomotor Target Predictor with a low-level Target-Conditioned Structured Dexterous Policy. Conditioned on egocentric RGB, proprioceptive and geometric history, the high-level predictor constructs structured hand-tool-object visuomotor embeddings and uses a horizon-conditioned transformer to generate a multi-step future target trajectory. Then, the low-level policy tracks them with a target-conditioned per-link transformer. This hierarchy decouples coarse future reference generation from fine-grained action control, and slow long-horizon semantic prediction from high-frequency execution. On OakInk2 bimanual tool-use tasks, DexFuture achieves 90% of the privileged-oracle performance, compared to 7% for a no-reference policy. DexFuture operates at 60 Hz, approximately 250 times faster than DexWM-style Cross-Entropy Method (CEM) planning with a future action-conditioned world model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。