arXiv:2607.08283cs.RO2026-07中稿 · the SemRob 2026 Wo…

让机器人理解任务进度,应对视觉相似却需不同动作的复杂操作。

TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning

论文配图:TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning
图 1 · 摘自论文原文
  • 用动态记忆追踪任务进展,关键动作前后主动更新信念。
  • 在多个数据集上成功率达98.75%,触觉任务诊断达75%成功率。
  • 适合需要记忆与阶段判断的复杂机械臂操作场景。

视觉-语言-动作(VLA)策略如π_{0.5}和OpenVLA在许多操作任务中表现良好,但通常为反应式:仅根据当前观测、指令和本体感觉预测下一步动作。这一假设在阶段依赖型操作中失效,因视觉相似状态可能因潜在任务进度和先前交互结果而需不同动作。我们提出时间条件记忆融合策略(TFP),一种轻量级记忆-动作框架,适用于VLA主干网络。TFP通过液态时间常数动力学维护每轮任务进度信念,并通过自适应调制将更新后的信念直接注入流匹配动作解码器。这使时间累积上下文主动塑造生成的动作片段,而非仅作为被动历史。使用3.3B参数模型,TFP在LIBERO上平均成功率从96.9%提升至98.75%,在LIBERO-plus上从91.4%提升至93.77%。在聚焦记忆的MIKASA ShellGameTouch诊断任务中,成功率达75.0%。机制分析显示,操作事件附近写入增益约为非事件期的6倍,隐藏状态干预表明信念对生成动作片段具有因果调控作用。结果表明,紧凑且事件敏感的记忆动态可显著提升在遮挡、视觉扰动和阶段依赖结构下的VLA性能。

原文摘要 · Abstract (English)

Vision--Language--Action (VLA) policies such as $π_{0.5}$ and OpenVLA perform well on many manipulation tasks, but they are often reactive: the next action is predicted from the current observation, instruction, and proprioceptive state. This assumption breaks down in stage-dependent manipulation, where visually similar states may require different actions depending on latent task progress and previous interaction outcomes. We argue that such tasks require not only memory, but dynamics-aware belief updates: the policy should preserve task progress during stable or occluded phases and revise its belief near contact, release, or subgoal transitions. We introduce Temporally Conditioned Memory-Fusion Policies (TFP), a lightweight memory-action framework for VLA backbones. TFP maintains an episode-local task-progress belief with Liquid Time-Constant dynamics and injects the updated belief directly into the flow-matching action decoder through adaptive modulation. This lets temporally accumulated context shape the generated action chunk, rather than serving only as passive history context. With a 3.3B-parameter model, TFP improves the average success rate from 96.9% to 98.75% on LIBERO and from 91.4% to 93.77% on LIBERO-plus. On the memory-focused MIKASA ShellGameTouch diagnostic, TFP achieves success up to 75.0%. Mechanistic analyses show that write-gain changes near manipulation events are about 6 times larger than far non-event phases, and hidden-state interventions show that the belief causally modulates generated action chunks. These results suggest that compact, event-sensitive memory dynamics can improve VLA policies under occlusion, visual perturbation, and stage-dependent task structure.

机器人控制记忆机制视觉-语言-动作任务进度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。