让视觉语言动作模型的内部特征可观察可控制,实现无需微调的实时行为调整。
Observing and Controlling Features in Vision-Language-Action Models
- 用线性分类器观测模型内部特征,揭示其可解释性
- 通过轻量级线性干预精准引导机器人输出,效果稳定可靠
- 适用于多种VLA架构,支持实时用户偏好对齐
视觉-语言-动作模型(VLAs)在具身智能方面取得显著进展。尽管其架构部分借鉴大语言模型(LLMs),但因多模态输入/输出及常混合使用Transformer与扩散头,复杂度更高,导致LLM中的机械可解释性洞察难以直接迁移。本文提出并分析两个核心概念:特征可观测性与特征可控性。首先研究表征空间中线性编码的特征,证明可通过线性分类器进行观测;随后,基于最优控制设计最小线性干预,精确调控内部表示,引导VLA输出至期望区域。实验在π₀.₅与OpenVLA两种VLA架构上通过仿真验证,表明目标明确、轻量化的干预能可靠引导机器人行为,同时保持闭环能力。结果表明,无需微调即可在线适配,实现实时对齐用户偏好与任务需求。
原文摘要 · Abstract (English)
Vision-Language-Action Models (VLAs) have shown remarkable progress towards embodied intelligence. While their architecture partially resembles that of Large Language Models (LLMs), VLAs exhibit higher complexity due to their multi-modal inputs/outputs and often hybrid nature of transformer and diffusion heads. This is part of the reason why insights from mechanistic interpretability in LLMs, which explain how the internal model representations relate to their output behavior, do not trivially transfer to VLA counterparts. In this work, we propose to close this gap by introducing and analyzing two main concepts: feature-observability and feature-controllability. In particular, we first study features that are linearly encoded in representation space, and show how they can be observed by means of a linear classifier. Then, we use a minimal linear intervention grounded in optimal control to accurately place internal representations and steer the VLA's output towards a desired region. Our results show that targeted, lightweight interventions can reliably steer a robot's behavior while preserving closed-loop capabilities. We demonstrate on different VLA architectures ($π_{0.5}$ and OpenVLA) through simulation experiments that VLAs possess interpretable internal structure amenable to online adaptation without fine-tuning, enabling real-time alignment with user preferences and task requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。