通过时间对齐记忆,用低频修正提升视觉语言动作模型的长程操作成功率。
Retrieve in Time, Correct in Frequency

- 按时间顺序对齐成功轨迹,精准定位应修正的动作片段。
- 在4个LIBERO任务中成功率从86.4%提升至88.4%,长程任务达68.6%。
- 无需参数更新,可在客户端CPU实时运行,单次推理后完成修正。
冻结的视觉-语言-动作(VLA)策略生成时序扩展的动作块,但长程操作仍易受执行误差累积和视觉混淆影响。成功推演提供有效纠错依据,但现有帧检索可能返回与进展错位的动作,直接重放或时域融合会破坏策略的反应式结构。本文提出训练无关的测试时修正框架RTCF,通过分离经验检索与动作转移部分,实现高效修正。渐进式记忆对齐(PMA)通过递增更新单调边界,将不断增长的视觉执行历史与完整成功轨迹进行因果对齐,无需阶段标签即可联合识别相关记忆及其当前对齐位置。从对齐动作块中,RTCF对运动通道施加系数截断的低频残差修正,而高频分量与夹爪决策仍继承自冻结策略。在四个LIBERO套件上,每条件2000个回合下,聚合成功率从86.4%提升至88.4%,LIBERO-Long提升至68.6%。这些改进无需参数更新、重复VLA推理或额外GPU资源:修正可在单次策略调用后于客户端CPU完成,每动作块中位延迟仅10.99毫秒。
原文摘要 · Abstract (English)
Frozen vision-language-action (VLA) policies generate temporally extended action chunks, but long-horizon manipulation remains vulnerable to accumulated execution error and visual aliasing across task stages. Successful rollouts provide useful corrective evidence, yet current frame retrieval can return progress-misaligned actions,while direct replay or time-domain fusion can overwrite the reactive structure of the policy proposal. We introduce Retrieve in Time, Correct in Frequency (RTCF), a training-free test-time correction framework that improves frozen VLA performance with low model-side overhead.RTCF separates which experience to retrieve from which part of its action to transfer. Progressive Memory Alignment (PMA) causally aligns the growing visual execution history with complete successful trajectories through incrementally updated monotonic frontiers, jointly identifying a relevant memory and the current aligned memory position without stage labels. From the aligned action chunk,RTCF transfers a coefficient-wise-clipped low-frequency residual on motion channels. Higher-frequency components and gripper decisions remain inherited from the frozen policy. Across four LIBERO suites and 2,000 episodes per condition, RTCF raises aggregate success from 86.4% to 88.4% and improves LIBERO-Long from 61.6% to 68.6%.These gains require no parameter updates, repeated VLA inference, or additional GPU resources: correction can be performed on the client CPU after a single policy invocation, and the median latencies sum to only 10.99 ms per action chunk
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。