无需训练即可完成复杂抓取与工具使用,靠多视角视觉理解指令
ZeroDex: Zero-Shot Long-Horizon Dexterous Manipulation via Multi-View 3D-Grounded VLM Reasoning

- 用多视角图像和视觉语言模型生成3D任务规划点
- 3D关键点精度提升,支持零样本抓取与工具操作
- 适合机器人长时序灵巧操作研究者
我们提出ZeroDex,一种零样本长时序灵巧操作框架,通过校准的多视角RGB图像将语言指令转化为可执行的3D任务计划。系统不依赖端到端策略训练,而是利用视觉语言模型(VLM)生成参考坐标系下的任务定位和原始级别的2D关键点,再通过多视角融合将其提升至3D。该过程结合视图内VLM定位的三角测量与参考视图射线投票,沿语义相机射线搜索跨邻近视图的几何一致性候选点。生成的3D关键点支持抓取放置与工具使用:对于工具使用,检索对应推断技能类别的物体中心原子动作,并将其存储的6D工具轨迹对齐至场景;对于灵巧操作,将提升的抓取关键点扩展为任务条件化的抓取可达区域,并通过机械臂-手部运动生成器生成可行的抓取-运动组合。真实世界实验表明,相比单视角RGB-D定位和微调的视觉语言代理基线,本方法在3D定位精度与执行可靠性上均有提升。我们进一步通过闭环状态验证与重规划实现长时序操作,成功在未见物体与新场景中完成零样本工具使用任务。
原文摘要 · Abstract (English)
We present ZeroDex, a zero-shot framework for long-horizon dexterous manipulation that grounds language instructions into executable 3D task plans from calibrated multi-view RGB images. Rather than training an end-to-end policy, our system uses a vision-language model (VLM) to produce reference-frame task grounding and primitive-level 2D keypoints, then lifts them into 3D via multi-view fusion. This lifting combines triangulation of view-wise VLM groundings with reference-view ray voting, which searches along a semantic camera ray for geometrically consistent candidates across neighboring views. The resulting 3D keypoints support both pick-and-place and tool-use: for tool-use, we retrieve an object-centric atomic action corresponding to the inferred skill category and align its stored 6D tool trajectory to the scene; for dexterous execution, we expand the lifted grasp keypoint into a task-conditioned grasp affordance region and generate feasible grasp-motion pairs with an arm-hand motion generator. Real-world experiments show improved 3D grounding accuracy and execution reliability over single-view RGB-D grounding and fine-tuned VLA baselines. We further demonstrate long-horizon manipulation through closed-loop status verification and replan, enabling zero-shot execution on unseen objects and tool-use tasks in novel scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。