arXiv:2606.00269cs.AI2026-06中稿 · CVPR

让视觉语言动作模型实时自适应调整控制,提升任务稳定性。

Closed-Loop Neural Activation Control in Vision-Language-Action Models

论文配图:Closed-Loop Neural Activation Control in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用动态反馈信号替代固定强度干预,实现闭环控制
  • 在四个LIBERO任务中成功提升概念调节稳定性和任务成功率
  • 无需重训练模型,适合需要平滑动作的机器人控制场景

视觉-语言-动作(VLA)模型可在推理时通过干预语义内部方向进行调控,但现有方法使用固定调节系数,属于开环操作。这在具身控制中表现不佳,因任务状态和概念误差随时间变化,常导致过度纠正、振荡和任务成功率下降,尤其影响速度与流畅性等时序行为。本文提出CTRL-STEER,一种闭环框架,将静态干预强度替换为随时间自适应的控制信号。核心思路是解耦表示与调控:不假设时序概念由单个神经元直接控制,而是沿运动对齐的残差方向施加干预,并由反馈控制器在线调整干预强度。该框架采用基于PID和强化学习的控制器。在微调后的OpenVLA策略上,于四个LIBERO任务套件上的实验表明,CTRL-STEER实现了更稳定的概念调控,且在调控精度与任务成功率之间取得更好权衡,无需修改或重新训练基础模型。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models can be steered at test time by intervening on semantically meaningful internal directions, but existing methods use a fixed steering coefficient, effectively operating in open loop. This is poorly suited to embodied control, where task state and concept error evolve over time, often causing overcorrection, oscillation, and reduced task success, especially for temporal behaviors such as speed and smoothness. We propose CTRL-STEER, a closed-loop framework that replaces static intervention strength with adaptive, time-varying control signals. The key idea is to decouple representation from regulation: rather than assuming temporal concepts are directly controlled by individual neurons, we steer along motion-aligned residual directions while a feedback controller adjusts intervention magnitude online. We instantiate this framework with both PID and reinforcement learning based controllers. Experiments with a fine-tuned OpenVLA policy on four LIBERO task suites show that CTRL-STEER achieves more stable concept regulation and a better steering-task success trade-off than fixed-coefficient baselines, without modifying or retraining the base model.

闭环控制具身智能模型调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。