arXiv:2510.05662cs.ROcs.CV2025-10被引 3

用语言和演示指导机器人精准操控透明物体,支持新物体长程操作。

DeLTa: Demonstration and Language-Guided Novel Transparent Object Manipulation

  • 结合深度估计、6D姿态与视觉语言规划,实现语言引导的透明物体操控。
  • 仅需一次演示即可泛化到新透明物体,无需类别先验或额外训练。
  • 适用于单臂眼在手上机器人,适合长程高精度任务,适合实际场景应用。

尽管透明物体在日常生活中广泛存在,但现有机器人操控研究仍局限于短时任务和基础抓取。虽然部分方法已部分解决此问题,但大多难以泛化到新物体,且缺乏对长程精确操控的支持。为此,我们提出 DeLTa(演示与语言引导的新透明物体操控框架),融合深度估计、6D 姿态估计与视觉语言规划,实现由自然语言指令引导的长程精确操控。该方法的关键优势在于仅需一次演示即可将 6D 轨迹泛化至新透明物体,无需类别先验或额外训练。此外,我们设计了一个任务规划器,优化视觉语言模型生成的计划以适应单臂眼在手上机器人的约束。全面评估表明,本方法显著优于现有透明物体操控方案,尤其在需要高精度的长程任务中表现突出。

原文摘要 · Abstract (English)

Despite the prevalence of transparent object interactions in human everyday life, transparent robotic manipulation research remains limited to short-horizon tasks and basic grasping capabilities. Although some methods have partially addressed these issues, most of them have limitations in generalization to novel objects and are insufficient for precise long-horizon robot manipulation. To address this limitation, we propose DeLTa (Demonstration and Language-Guided Novel Transparent Object Manipulation), a novel framework that integrates depth estimation, 6D pose estimation, and vision-language planning for precise long-horizon manipulation of transparent objects guided by natural language task instructions. A key advantage of our method is its single-demonstration approach, which generalizes 6D trajectories to novel transparent objects without requiring category-level priors or additional training. Additionally, we present a task planner that refines the VLM-generated plan to account for the constraints of a single-arm, eye-in-hand robot for long-horizon object manipulation tasks. Through comprehensive evaluation, we demonstrate that our method significantly outperforms existing transparent object manipulation approaches, particularly in long-horizon scenarios requiring precise manipulation capabilities. Project page: https://sites.google.com/view/DeLTa25/

透明物体语言引导长程操控6D姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。