让视觉预测模型在测试时自我优化,提升对异常情况的适应能力。
Test-Time Training for Visual Foresight Vision-Language-Action Models

- 测试时利用预测图像与真实观测构成监督信号,持续更新模型。
- 在多个分布外场景下,性能下降幅度减少超过50%。
- 无需修改结构或额外模块,适合部署在实际机器人系统中。
视觉预见型视觉-语言-动作模型(VF-VLA)近年来表现出色,但其架构对分布外(OOD)变化极为敏感。由于动作质量依赖于未来视觉信息的预测准确性,分布外条件会同时影响预测和执行阶段。为此,本文提出测试时训练方法 $T^3$VF,基于预测未来图像与其后续真实观测之间天然形成的监督对进行自适应更新。为应对测试时更新带来的实际挑战,引入动态更新过滤机制。实验表明,$T^3$VF在仅增加少量推理开销的前提下,有效缓解了VF-VLA的分布外脆弱性,且无需任何结构改动或附加模块。
原文摘要 · Abstract (English)
Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inherent design of VF-VLA makes it particularly vulnerable to out-of-distribution (OOD) shifts. Because the quality of action directly depends on the accuracy of the predicted future visual information, OOD conditions affect both stages at once. To address this vulnerability, we propose Test-Time Training Visual Foresight VLA ($T^3$VF), a test-time training approach motivated by the observation that the predicted future image and its subsequent observation form a natural supervision pair. To further address the practical challenges that arise from indiscriminate test-time updates, we introduce an adaptive update filtering mechanism. Empirically, $T^3$VF mitigates the OOD vulnerability of VF-VLA at a modest additional inference cost, without requiring any architectural modification or auxiliary modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。