arXiv:2604.05614cs.RO2026-04

让机器人说话与动作对得上,提升人机协作透明度。

Grounding Hierarchical Vision-Language-Action Models Through Explicit Language-Action Alignment

  • 用对比模型衡量语言与动作轨迹的匹配度,实现显式对齐。
  • 在LanguageTable数据集上达到接近全监督微调的性能。
  • 减少昂贵标注依赖,适合需要高可解释性的机器人应用。

实现机器人透明性是促进人机协作的关键。为使机器人具备透明性,其自然语言表达必须与动作一致,并明确基于任务和环境。现有分层视觉-语言-动作(VLA)模型虽能生成语言(如思维链)和低级动作,但训练时未显式对齐多模态信息。为此,我们提出一种新训练框架,将分层VLA的子任务描述显式地与视觉观测和动作空间对齐。该框架采用对比模型评估生成语言与对应动作轨迹之间的对齐程度,通过离线偏好学习直接排序不同语言-轨迹组合,从而优化模型的语义接地能力。我们在LanguageTable数据集(人类语言标注轨迹的基准数据集)上验证了方法,揭示了多模态接地表示的关键特性,同时建立了一个强基线,性能媲美全监督微调,且显著降低对高成本标注数据的需求。

原文摘要 · Abstract (English)

Achieving robot transparency is a critical step toward effective human-robot collaboration. To be transparent, a robot's natural language communication must be consistent with its actions and explicitly grounded in the task and environment. Existing hierarchical Vision-Language-Action (VLA) models can generate language (e.g., through chain-of-thought) and low-level actions. However, current work does not consider explicit alignment between these modalities during training. To address this crucial gap, we propose a novel training framework that explicitly grounds hierarchical VLA sub-task descriptions with respect to the visual observation and action space. Our framework uses a contrastive model to assess the alignment between generated language and corresponding action trajectories. This contrastive model enables direct ranking of different language-trajectory pairs based on their alignment, allowing us to refine the grounding of our hierarchical VLA through offline preference learning. We apply our framework to the LanguageTable dataset, a benchmark dataset of human language-annotated trajectories, and provide critical insights into multimodal grounding representations, all while establishing a strong baseline that achieves performance comparable to fully supervised fine-tuning and minimizing the need for costly data annotations.

机器人多模态语言对齐可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。