arXiv:2608.09448cs.ROcs.CV2026-08

通过预测未来视觉结果,让视觉语言动作模型更可靠地在线适应。

VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction

论文配图:VANE: Reliable Test-Time Training for Vision-Language-Action Models via Future Visual Representation Prediction
图 1 · 摘自论文原文
  • 基于当前视觉语言上下文和未来视觉反馈进行选择性更新
  • 在WidowX上成功率提升3.2个百分点,效果可逆且可控
  • 适合需要安全在线调整的机器人操作场景

测试时训练(TTT)为从无标签部署流中轻量级适配视觉-语言-动作(VLA)策略提供了可能,但在闭环操控中仍难以可靠使用。共享的适应空间可能导致不兼容的任务修正混淆,而在线更新可能在行动后果尚未可知时就影响后续动作。本文提出一种可靠的VLA策略测试时训练框架VANE。VANE将提示适配条件限定在当前视觉-语言上下文,并从执行动作的未来视觉后果中学习。候选更新被隔离于主策略之外,根据后续观测进行评估,仅在有未来证据支持时才提交,确保适应过程选择性强且可逆。在SimplerEnv WidowX上,VANE相比对应基线平均成功率提升3.2个百分点。谷歌机器人实验进一步表明,部署阶段的收益具有任务与本体依赖性。这些结果共同展示了在交互过程中采用受限、基于证据的VLA策略适应方法的有效性。

原文摘要 · Abstract (English)

Test-time training (TTT) offers a lightweight way to adapt vision--language--action (VLA) policies from unlabeled deployment streams, but it remains difficult to use reliably in closed-loop manipulation. A shared adaptation space can mix incompatible task corrections, while an online update can alter subsequent actions before its consequences are known. We introduce a reliable TTT framework for VLA policies (VANE). VANE conditions prompt adaptation on the current vision--language context and learns from the future visual consequences of executed actions. Candidate updates are isolated from the live policy, evaluated on subsequent observations, and committed only when supported by future evidence, making adaptation selective and reversible. On SimplerEnv WidowX, VANE improves average success by $3.2$ percentage points over the corresponding TTT baseline. Results on Google Robot further show that deployment-time gains remain task- and embodiment-dependent. Together, these results demonstrate a constrained, evidence-based approach to adapting VLA policies during interaction.

在线适应机器人控制视觉语言动作可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。