用信息论融合视觉语言动作,让双臂机器人从人类视频中学会复杂操作。
Information-Theoretic Graph Fusion with Vision-Language-Action Model for Policy Reasoning and Dual Robotic Control
- 基于香农信息量识别高相关性的手与物体,构建时序场景图
- 实现95%图准确率、93%子任务分割,支持生成可读任务策略
- 无需几何计算即可优化双手抓取分配,适合复杂空间操作
由于依赖低级轨迹模仿,从人类视频学习灵巧技能仍具挑战性,难以泛化到不同物体类型、空间布局和机械臂配置。本文提出图融合视觉-语言-动作(GF-VLA)框架,使双臂机器人能直接从RGB和深度人类示范中进行任务级推理与执行。该框架首先提取基于香农信息的线索,识别任务相关性最高的手与物体,再将其编码为包含手-物及物-物交互的时序场景图。这些图与语言条件化的Transformer融合,生成层次化行为树和可解释的笛卡尔运动指令。为提升双臂执行效率,进一步引入跨手选择策略,无需显式几何推理即可推断最优夹持器分配。在四类结构化双臂积木组装任务中评估,包括符号形状构建与空间泛化。实验显示,信息论场景表示达到超过95%的图准确率和93%的子任务分割率,支持大模型规划器生成可靠且可读的任务策略。在双臂机器人上执行时,抓取成功率94%,放置准确率89%,整体任务成功率90%,涵盖堆叠、字母构建与几何重构场景,表现出强泛化能力与鲁棒性。
原文摘要 · Abstract (English)
Teaching robots dexterous skills from human videos remains challenging due to the reliance on low-level trajectory imitation, which fails to generalize across object types, spatial layouts, and manipulator configurations. We propose Graph-Fused Vision-Language-Action (GF-VLA), a framework that enables dual-arm robotic systems to perform task-level reasoning and execution directly from RGB and Depth human demonstrations. GF-VLA first extracts Shannon-information-based cues to identify hands and objects with the highest task relevance, then encodes these cues into temporally ordered scene graphs that capture both hand-object and object-object interactions. These graphs are fused with a language-conditioned transformer that generates hierarchical behavior trees and interpretable Cartesian motion commands. To improve execution efficiency in bimanual settings, we further introduce a cross-hand selection policy that infers optimal gripper assignment without explicit geometric reasoning. We evaluate GF-VLA on four structured dual-arm block assembly tasks involving symbolic shape construction and spatial generalization. Experimental results show that the information-theoretic scene representation achieves over 95 percent graph accuracy and 93 percent subtask segmentation, supporting the LLM planner in generating reliable and human-readable task policies. When executed by the dual-arm robot, these policies yield 94 percent grasp success, 89 percent placement accuracy, and 90 percent overall task success across stacking, letter-building, and geometric reconfiguration scenarios, demonstrating strong generalization and robustness across diverse spatial and semantic variations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。