arXiv:2605.09537cs.RO2026-05中稿 · ICML被引 1

用信号噪声比动态触发搜索,解决机器人长任务中的指令漂移问题

Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning

论文配图:Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning
图 1 · 摘自论文原文
  • 通过功率分布增强全局轨迹概率,实现推理时的前瞻搜索
  • 在多个基准上超越OpenVLA和TACO,无需参数更新
  • 基于信噪比的元认知机制,智能切换快速决策与慢速搜索

尽管视觉-语言-动作(VLA)模型在机器人控制中进展迅速,长时序任务中的指令漂移仍是持续存在的故障模式。本文重新定义该现象,认为指令漂移本质上是系统性采样误差:局部贪婪采样易陷入‘负关键窗口’——高局部概率但不可逆的局部最优,切断全局成功路径。为此,提出无训练的推理时计算框架上下文感知功率采样(CAPS),利用功率分布强化全局轨迹概率,实现对模型条件生成轨迹分布的前瞻搜索。进一步引入基于信噪比(SNR)的元认知控制机制,仅在检测到漂移风险时触发自适应MCMC搜索,实现从‘直觉快思考’到‘理性慢搜’的动态切换。在RoboTwin、Simpler-WindowX和Libero-long基准上的实验表明,CAPS在不更新参数的情况下显著优于OpenVLA和TACO等强基线,验证了自适应推理时计算在具身控制长时序鲁棒性提升中的有效性。

原文摘要 · Abstract (English)

Despite rapid progress in Vision-Language-Action (VLA) models for robotic control, instruction drift remains a persistent failure mode in long-horizon tasks. This paper reconceptualizes this phenomenon, positing that instruction drift is fundamentally a systematic sampling error: local greedy sampling is prone to collapsing into "Negative Pivotal Windows"--irreversible local optima with high local probability that sever global success pathways. To address this, we propose Context-Aware Power Sampling (CAPS), a training-free inference-time computation framework. CAPS leverages power distributions to sharpen global trajectory probabilities, enabling lookahead search over the model's conditional generative trajectory distribution. Furthermore, we introduce a metacognitive control mechanism based on Signal-to-Noise Ratio (SNR). This mechanism triggers adaptive MCMC search solely when drift risk is detected, enabling a dynamic transition from "intuitive fast thinking" to "rational slow search." Experiments on RoboTwin, Simpler-WindowX, and Libero-long benchmarks show that CAPS achieves substantial improvements over strong baselines, including OpenVLA and TACO, without parameter updates. These results support the effectiveness of adaptive inference-time computation for improving long-horizon robustness in embodied control.

机器人规划采样优化元认知长时序控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。