arXiv:2606.30318cs.RO2026-06被引 2

Chronos让机器人记住操作历史,实现更精准的长时序动作控制。

Chronos: A Physics-Informed Full-History Framework for Non-Markovian Long-Horizon Manipulation

论文配图:Chronos: A Physics-Informed Full-History Framework for Non-Markovian Long-Horizon Manipulation
图 1 · 摘自论文原文
  • 将历史观测作为策略动态的隐状态,实现因果时间对齐
  • 在真实场景中达78%成功率,内存任务上72%优于基线
  • 参数量少10倍,却显著超越现有方法,适合复杂操控

通用机器人策略应建模为动力系统,但多数视觉语言模型与生成式模仿策略仍依赖当前观测或短窗口。这种马尔可夫简化在依赖记忆的操作中失效:相同观测可能因历史不同而需不同动作。本文提出Chronos,一种物理信息驱动的全历史非马尔可夫长时序操控框架。核心思想是将观测历史从辅助上下文提升为策略动态的隐状态。每一步物理控制中,Chronos通过融合观测与本体感知生成一个状态代表性标记,使标记序列与物理时间一一对应。选择性状态空间模型传播此因果历史状态,并通过隐式最大似然估计(IMLE)生成多模态粗动作先验。该先验再经二阶薛定谔启发桥接模型优化,预测加速度场,从而生成更平滑、更符合物理规律的机器人运动。在16个仿真任务和4个真实实验中评估,涵盖精密插入、通用操控及记忆依赖长时序控制。在RMBench上,任务成功需记忆阶段,Chronos平均成功率73.6%,较马尔可夫基线pi0.5提升62.4个百分点(相对增益6.6倍),仅用1/10参数;较内存模型Mem-0高22.8点,且参数量超30倍少。真实双臂实验中,单摄像头下四任务平均成功率78%,其中三个记忆依赖任务达72%;而pi0.5整体仅7%,记忆任务为0%。结果表明,历史不应视为辅助上下文,而应作为操控策略的隐状态。

原文摘要 · Abstract (English)

General-purpose robot policies should be modeled as dynamical systems, yet many VLA and generative imitation policies still rely on present observations or short windows. This Markovian shortcut fails in memory-dependent manipulation: identical observations can demand different actions after different histories. We present Chronos, a physics-informed full-history framework for non-Markovian long-horizon manipulation. The key idea is to elevate observation history from auxiliary context to the latent state of the policy dynamics. At each physical control step, Chronos forms one state-representative token by fusing observation and proprioception, so the token sequence is aligned one-to-one with physical time. A selective state space model propagates this causal historical state, which conditions a multimodal coarse action prior through implicit maximum likelihood estimation (IMLE). This prior is then refined by a second-order Schrodinger-inspired bridge that predicts acceleration fields, yielding smoother and more physically grounded robot motion. Across 16 simulated tasks and 4 real-world experiments, Chronos is evaluated on precision insertion, general manipulation, and memory-dependent long-horizon control. On RMBench, where success requires remembering task phase, Chronos achieves 73.6% average success, outperforming Markovian VLA baseline pi0.5 by +62.4 percentage points, a 6.6x relative gain, while using 10x fewer parameters. It also surpasses the memory VLA Mem-0 by 22.8 points while using over 30x fewer parameters. In real-world dual-arm experiments using a single RGB camera, Chronos achieves 78% average success over four tasks, including 72% on the three memory-dependent tasks, whereas pi0.5 achieves 7% overall and 0% on the memory-dependent subset. These results suggest that history should not be treated as auxiliary context, but as the latent state of the manipulation policy.

机器人操控长时序控制物理模型历史记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。