arXiv:2509.07957cs.RO2025-09被引 1

用视觉语言动作图融合,让双臂机器人从人类示范中学会灵活抓取组装。

Graph-Fused Vision-Language-Action for Policy Reasoning in Multi-Arm Robotic Manipulation

  • 构建时空场景图提取关键交互,结合语言模型生成分层动作树
  • 在4个任务上实现95%图准确率、93%子任务分割,部署后任务成功率90%
  • 无需几何建模即可自动分配双手,适合复杂多臂操作场景

从人类视频示范中获取灵巧机器人技能仍面临挑战,主要源于传统方法依赖低层轨迹复制,难以在不同物体、空间布局和机械臂配置下泛化。为此,我们提出图融合视觉-语言-动作(GF-VLA)框架,使双臂机器人能直接从RGB-D人类示范中进行任务级推理与执行。该框架采用信息论方法提取任务相关线索,聚焦关键手物及物物交互,将其结构化为时序有序的场景图,并与语言条件变压器融合,生成层次化行为树和可解释的笛卡尔运动原语。为提升双臂执行效率,提出跨臂分配策略,无需显式几何建模即可自主确定夹持器分配。我们在四个涉及符号结构构建与空间泛化的双臂积木装配基准上验证了该方法。实证结果表明,所提表示在图准确率超过95%、子任务分割率达93%的基础上,语言-动作规划器可生成鲁棒且可解释的任务策略。在双臂机器人部署中,策略达到94%抓取可靠性、89%放置精度和90%整体任务成功率,展现出在多样空间与语义变化下的强泛化能力与鲁棒性。

原文摘要 · Abstract (English)

Acquiring dexterous robotic skills from human video demonstrations remains a significant challenge, largely due to conventional reliance on low-level trajectory replication, which often fails to generalize across varying objects, spatial layouts, and manipulator configurations. To address this limitation, we introduce Graph-Fused Vision-Language-Action (GF-VLA), a unified framework that enables dual-arm robotic systems to perform task-level reasoning and execution directly from RGB-D human demonstrations. GF-VLA employs an information-theoretic approach to extract task-relevant cues, selectively highlighting critical hand-object and object-object interactions. These cues are structured into temporally ordered scene graphs, which are subsequently integrated with a language-conditioned transformer to produce hierarchical behavior trees and interpretable Cartesian motion primitives. To enhance efficiency in bimanual execution, we propose a cross-arm allocation strategy that autonomously determines gripper assignment without requiring explicit geometric modeling. We validate GF-VLA on four dual-arm block assembly benchmarks involving symbolic structure construction and spatial generalization. Empirical results demonstrate that the proposed representation achieves over 95% graph accuracy and 93% subtask segmentation, enabling the language-action planner to generate robust, interpretable task policies. When deployed on a dual-arm robot, these policies attain 94% grasp reliability, 89% placement accuracy, and 90% overall task success across stacking, letter-formation, and geometric reconfiguration tasks, evidencing strong generalization and robustness under diverse spatial and semantic variations.

双臂机器人视觉语言动作规划场景图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。