arXiv:2506.17328cs.RO2025-06被引 1

用视觉语言模型规划双臂清洁,实现复杂杂物精准抓取

Reflective VLM Planning for Dual-Arm Desktop Cleaning: Bridging Open-Vocabulary Perception and Precise Manipulation

  • 通过反思式VLM生成并优化双臂操作序列
  • 模拟环境下任务完成率达87.2%,较单臂提升36.2%
  • 适合需要多模态感知与精确操控的桌面级机器人任务

桌面清洁需具备开放词汇识别与精确操作能力以应对异构垃圾。我们提出一种分层框架,将反思式视觉语言模型(VLM)规划与双臂执行结合,基于结构化场景表示实现端到端控制。Grounded-SAM2 实现开放词汇检测,记忆增强型 VLM 负责生成、评估并修正操作序列,再转换为五种基本动作的参数化轨迹,由协同工作的 Franka 机械臂执行。在模拟场景中,系统任务完成率达 87.2%,相比静态 VLM 提升 28.8%,较单臂基线提升 36.2%。结构化记忆集成对实现鲁棒、可泛化的操作至关重要,同时保持实时控制性能。

原文摘要 · Abstract (English)

Desktop cleaning demands open-vocabulary recognition and precise manipulation for heterogeneous debris. We propose a hierarchical framework integrating reflective Vision-Language Model (VLM) planning with dual-arm execution via structured scene representation. Grounded-SAM2 facilitates open-vocabulary detection, while a memory-augmented VLM generates, critiques, and revises manipulation sequences. These sequences are converted into parametric trajectories for five primitives executed by coordinated Franka arms. Evaluated in simulated scenarios, our system achieving 87.2% task completion, a 28.8% improvement over static VLM and 36.2% over single-arm baselines. Structured memory integration proves crucial for robust, generalizable manipulation while maintaining real-time control performance.

双臂机器人视觉语言模型桌面清洁路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。