通过限制视觉信息输入,让模型更专注关键细节,提升机器人任务泛化能力。
See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

- 用精炼的语言标注和视觉证据预算,引导模型聚焦关键信息。
- 在真实机器人上,子任务成功率从54.2%提升至79.0%。
- 适合研究视觉-语言-动作模型泛化与机器人控制的开发者。
视觉-语言-动作(VLA)模型的泛化能力仍受制于干扰、外观变化及语义相似任务:策略需从粗略指令中推断局部执行细节,同时判断图像中哪些部分对控制重要。我们提出S2(See Less, Specify More)框架,通过更清晰的接口训练执行器以提升泛化性能。'Specify More'保留原始指令作为稳定高层目标,将每条轨迹重标注为细粒度的轨迹与子任务级语言,消除当前执行模式的歧义。'See Less'则施加显式的视觉证据预算,训练执行器仅基于任务充分的视觉证据行动,无需区域或掩码标注。该接口使执行器能在不依赖干扰视觉块或自行解决可避免歧义的情况下,遵循详细指导,且兼容现成的VLM规划器,通过上下文学习实现。在主评估设置中,S2通过改变执行器的学习问题提升整体泛化:粗略指令引发可避免的监督混淆,保持目标的局部指导优于指令替换,显式证据预算减少对广泛视觉上下文的依赖,不仅提升效率。在八项真实机器人任务(TX-G2 和 HSR)上,相比 pi0.5,S2 将平均子任务成功率从54.2%提升至79.0%。结果表明,当执行器被训练为基于信息性局部指导和任务充分的视觉证据行动时,而非从弱监督中恢复二者,其泛化能力显著增强。
原文摘要 · Abstract (English)
Generalization remains a central bottleneck for vision-language-action (VLA) models: under distractors, appearance shifts, and semantically similar tasks, the policy must often infer local execution details from coarse instructions while also deciding which parts of the image matter for control. We present S2 (See Less, Specify More), a framework for improving VLA generalization by training the executor under a cleaner interface. Specify More preserves the original instruction as a stable high-level goal while relabeling each trajectory into refined trajectory- and subtask-level language that disambiguates the current execution mode. Unlike native attention, See Less imposes an explicit visual evidence budget, training the executor to act from task-sufficient evidence rather than unconstrained visual context, without any region or mask annotation. This interface lets the executor follow detailed guidance without relying on distracting visual patches or resolving avoidable ambiguity on its own, and it remains compatible with off-the-shelf VLM planners through in-context learning. Across our main evaluation settings, S2 improves overall generalization metrics by changing the executor's learning problem: coarse instructions induce avoidable supervision aliasing, goal-preserving local guidance outperforms instruction replacement in our main ablations, and explicit evidence budgeting reduces dependence on broad visual context beyond efficiency considerations. Across eight real-robot tasks on TX-G2 (an AgiBot G2-compatible variant) and HSR, S2 raises mean subtask success from 54.2% to 79.0% over pi0.5. Together, these results suggest that VLA generalization improves when the executor is trained to act from informative local guidance and task-sufficient visual evidence, rather than recovering both from weak supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。