通过评估可用潜力与外部补足性,智能选择测试时增强策略。
Reuse Before You Retrieve: Diagnosing Headroom and Complementarity for Test-Time Augmentation of Embodied Multimodal Policies
- 用可恢复潜力和检索互补性量化测试时增强需求
- 在LIBERO上最高提升21.0%成功率,与潜力水平高度匹配
- 适用于机器人任务部署,尤其在观测退化时仍有效
冻结的视觉-语言-动作(VLA)策略通过采样额外行为或引入外部示范,在测试时持续优化。然而,缺乏指导决定应采用何种干预。额外采样仅在策略已有更好行为且可识别时有效;而检索则在策略无法可靠表示相关动作先验时更优。本文通过可恢复潜力和检索互补性两个可测量因素,分析这一决策:前者表征已有可用行为量,后者判断外部动作先验是否填补了可测空白。在可重试或并行执行环境下评估基于回合的重试选择器,结合多个冻结VLA模型与环境。该选择器在所有测试的VLA骨干网络上均显著恢复潜在能力,最大成功率达21.0个百分点,且与可恢复潜力高度一致。其效果可迁移至不同机器人与模拟器,并在观测质量下降时依然有效。自回归OpenVLA实验进一步揭示了可用潜力与候选回溯排序能力的区别。检索表现不同,对动作先验差距最大的策略提升最显著,且与选择器结合后带来额外增益。结果为测试时增强机会提供了实证依据,区分了来自冻结策略内部的可恢复能力与需外部引入的行为先验。
原文摘要 · Abstract (English)
Frozen vision-language-action (VLA) policies are increasingly improved at test time by sampling additional policy behaviors or introducing external demonstrations. Yet there is little guidance for deciding which intervention a deployed policy actually needs. Additional sampling is useful only when better behavior already exists within the policy's stochastic rollouts and can be identified, whereas retrieval is most useful when the relevant action prior is not reliably represented by the policy. We study this decision through two measurable factors, recoverable headroom and retrieval complementarity, which characterize how much useful behavior is already available to recover and whether an external action prior fills a measurable gap. We evaluate an episode-level retry selector under retryable or parallel execution, together with retrieval across multiple frozen VLA policies and environments. The selector consistently recovers substantial latent capability across all tested VLA backbones on LIBERO, with gains of up to 21.0 success-rate points that closely track recoverable headroom. It also transfers to a different robot and simulator and remains effective under degraded observations, while experiments with autoregressive OpenVLA illustrate the distinction between available headroom and the ability to rank candidate rollouts. Retrieval behaves differently, improving the policy with the largest measured action-prior gap and providing further gains when combined with selection. Together, these results provide an empirical basis for characterizing test-time augmentation opportunities by separating capability that can be recovered from the frozen policy from behavioral priors that may need to be introduced externally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。