arXiv:2607.14280cs.ROcs.LG2026-07

提出新方法实现对视觉语言动作模型的精细行为控制

DiMaS: Distribution Matching for Steering Vision-Language-Action Models

  • 通过分布匹配替代线性方向移动,实现更灵活的表示调控
  • 在两种先进视觉语言动作模型上验证了有效的行为控制能力
  • 解释了传统方法失效原因,适合研究机器人决策与可解释性者

基于流匹配的视觉-语言-动作(VLA)模型已成为机器人操作的强大策略,但精细行为控制——即通过干预内部表示来调控机器人执行任务的方式——仍缺乏深入探索。现有表示调制方法通常依赖线性方向编码行为特征,但我们发现其在视觉运动场景中表现不佳。为此,本文提出DiMaS:一种专为流匹配式VLA设计的分布匹配调制策略,通过在表示分布间进行迁移而非沿固定方向移动,实现了对行为的有效控制。我们在两个先进VLA模型上验证了该方法的有效性,并系统分析了其在任务差异逐渐增大时的泛化能力,揭示了行为控制转移的边界。进一步分析动作专家的表示结构表明:虽然行为特征可线性解码,却无法通过线性方式调制,这解释了传统方法失效的原因,也支撑了分布匹配设计的合理性。代码已公开于https://github.com/pegah-kh/dimas,更多结果与视频见https://pegah-kh.github.io/dimas/

原文摘要 · Abstract (English)

Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a robot performs a task by intervening on its internal representations. Representation steering is a well-established interpretability tool for language and vision-language models, where behavioral features are typically encoded as linear directions, but we show that these classic methods fall short in VLAs. We propose DiMaS, a Distribution-Matching Steering strategy tailored to flow-matching VLAs, which transports between representation distributions rather than shifting along a fixed direction, and show that it effectively controls behavior across two state-of-the-art VLAs. We further examine the generalizability of this strategy as the tasks it is learned from and evaluated on grow increasingly dissimilar, characterizing where behavioral control transfers and where it weakens. Finally, through an analysis of the representation structure of the action expert, we explain why classical linear steering falls short in the visuomotor setting: behavioral features are linearly decodable but not linearly steerable, which motivates the distribution-matching design of DiMaS. Our code is publicly available at https://github.com/pegah-kh/dimas, with additional results and videos at https://pegah-kh.github.io/dimas/

视觉语言动作行为控制表示调制机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。