arXiv:2508.05342cs.ROcs.AI2025-08中稿 · Information Fusion被引 54

用信息论融合视觉语言动作,让双臂机器人从人类视频中学会复杂操作。

Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control

  • 基于香农信息量识别高相关性的手与物体,构建时序场景图
  • 实现95%图准确率、93%子任务分割,支持生成可读任务策略
  • 无需几何计算即可优化双手抓取分配,适合复杂空间操作

由于依赖低级轨迹模仿,从人类视频学习灵巧技能仍具挑战性,难以泛化到不同物体类型、空间布局和机械臂配置。本文提出图融合视觉-语言-动作(GF-VLA)框架,使双臂机器人能直接从RGB和深度人类示范中进行任务级推理与执行。该框架首先提取基于香农信息的线索,识别任务相关性最高的手与物体,再将其编码为包含手-物及物-物交互的时序场景图。这些图与语言条件化的Transformer融合,生成层次化行为树和可解释的笛卡尔运动指令。为提升双臂执行效率,进一步引入跨手选择策略,无需显式几何推理即可推断最优夹持器分配。在四类结构化双臂积木组装任务中评估,包括符号形状构建与空间泛化。实验显示,信息论场景表示达到超过95%的图准确率和93%的子任务分割率,支持大模型规划器生成可靠且可读的任务策略。在双臂机器人上执行时,抓取成功率94%,放置准确率89%,整体任务成功率90%,涵盖堆叠、字母构建与几何重构场景,表现出强泛化能力与鲁棒性。

原文摘要 · Abstract (English)

Teaching robots dexterous skills from human videos remains challenging due to the reliance on low-level trajectory imitation, which fails to generalize across object types, spatial layouts, and manipulator configurations. We propose Graph-Fused Vision-Language-Action (GF-VLA), a framework that enables dual-arm robotic systems to perform task-level reasoning and execution directly from RGB and Depth human demonstrations. GF-VLA first extracts Shannon-information-based cues to identify hands and objects with the highest task relevance, then encodes these cues into temporally ordered scene graphs that capture both hand-object and object-object interactions. These graphs are fused with a language-conditioned transformer that generates hierarchical behavior trees and interpretable Cartesian motion commands. To improve execution efficiency in bimanual settings, we further introduce a cross-hand selection policy that infers optimal gripper assignment without explicit geometric reasoning. We evaluate GF-VLA on four structured dual-arm block assembly tasks involving symbolic shape construction and spatial generalization. Experimental results show that the information-theoretic scene representation achieves over 95 percent graph accuracy and 93 percent subtask segmentation, supporting the LLM planner in generating reliable and human-readable task policies. When executed by the dual-arm robot, these policies yield 94 percent grasp success, 89 percent placement accuracy, and 90 percent overall task success across stacking, letter-building, and geometric reconfiguration scenarios, demonstrating strong generalization and robustness across diverse spatial and semantic variations.

双臂控制视觉语言图神经网络机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。