让机器人在不确定时既重看又重做,不需额外训练
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models
- 根据自身不确定性动态调整视觉感知和动作决策
- 单次前向传播实现感知与行动的联合优化,提升鲁棒性
- 适合需要实时响应的机器人控制场景,无需额外模型
视觉-语言-动作(VLA)模型已成为通用机器人控制的有前景范式,测试时缩放(TTS)方法被关注以增强训练外的鲁棒性。然而现有TTS方法需额外训练、验证器及多次前向传播,难以部署;且仅干预动作解码,固定视觉表征,在感知模糊时效果有限。为此,我们提出SCALE,一种基于自不确定性调节的简单推理策略,受主动推断理论中不确定性驱动探索启发,无需额外训练、验证器,仅需一次前向传播。在高不确定性下,同时扩大感知与动作的探索范围;低不确定性时则聚焦利用,实现跨条件自适应执行。在模拟与真实世界基准上的实验表明,SCALE可提升现有先进VLA模型性能,并优于现有TTS方法,同时保持单次前向传播效率。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic control, with test-time scaling (TTS) gaining attention to enhance robustness beyond training. However, existing TTS methods for VLAs require additional training, verifiers, and multiple forward passes, making them impractical for deployment. Moreover, they intervene only at action decoding while keeping visual representations fixed-insufficient under perceptual ambiguity, where reconsidering how to perceive is as important as deciding what to do. To address these limitations, we propose SCALE, a simple inference strategy that jointly modulates visual perception and action based on 'self-uncertainty', inspired by uncertainty-driven exploration in Active Inference theory-requiring no additional training, no verifier, and only a single forward pass. SCALE broadens exploration in both perception and action under high uncertainty, while focusing on exploitation when confident-enabling adaptive execution across varying conditions. Experiments on simulated and real-world benchmarks demonstrate that SCALE improves state-of-the-art VLAs and outperforms existing TTS methods while maintaining single-pass efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。