让机器人学会用陌生工具完成相同功能,靠的是关键点轨迹推理。
FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning

- 用关键点轨迹表示工具功能,分离语义理解与动作执行
- 在七种未见过的工具上成功率超现有方法2倍以上
- 适合需要跨工具泛化的机器人操控研究者
人类能轻松将书本、石头或鞋子当作锤子使用,但机器人在训练过的特定工具上表现良好,面对新工具却无法迁移功能——这一差距被我们定义为功能性泛化。这些工具虽视觉相似、功能意图一致,但在动作空间中却需完全不同的运动模式。为此,我们探索了包括可操作性图像、人类视频提示和2D关键点轨迹在内的中间表征,发现关键点轨迹在功能表达力与动作可实现性之间达到最佳平衡。基于此,我们提出两阶段策略FORGE:先从无动作数据中预测通用的关键点轨迹,再通过少量示范将其转化为机器人动作。在包含七种工具的敲击任务基准测试中,FORGE在模拟环境与真实世界中均显著优于现有方法,在未见工具上的平均成功率提升超过2倍。
原文摘要 · Abstract (English)
While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to novel ones -- a gap we formalize as functional generalization. Such tools share a common functional intent that is visually recognizable, yet this perceptual similarity does not carry over to action space, where each tool demands an entirely different motor pattern. To bridge this gap, we explore intermediate representations including affordance images, human video prompts, and 2D keypoint trajectories, finding that keypoint trajectories best balance functional expressiveness and action groundability. Building on this, we propose FunctiOnal Reasoning and Grounded Execution (FORGE), a two-stage policy that decouples functional reasoning from action execution: predicting generalizable keypoint trajectories from action-free data, then grounding them into robot actions with limited demonstrations. On a seven-tool hitting-function benchmark, FORGE consistently outperforms state-of-the-art methods on unseen tools in both simulation and the real world, achieving over 2X improvement in average success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。