arXiv:2608.30883cs.RO2026-08

用未来物理信息训练机器人,让其在看不见时也能稳走

SleepWalking: Privileged Representation Shaping for End-to-End Blind Locomotion in Legged Robots

论文配图:SleepWalking: Privileged Representation Shaping for End-to-End Blind Locomotion in Legged Robots
图 1 · 摘自论文原文
  • 用未来物理状态重建来引导记忆保留,不改推理结构
  • 比最强基线高15%地形等级,推理能耗降44.4%
  • 适合想提升复杂环境行走能力的机器人研究者

部分可观测的运动控制要求策略在任务相关状态未被即时观测完全揭示时仍能行动。现有方法通常通过显式估计缺失物理变量或使用结构化架构处理扩展观测历史来应对。本文提出不同视角:部分可观测性本质是信息保留问题。关键不在于信息如何进入网络,而在于策略内部状态能否保留它。基于此,我们提出用于机器人行走的睡眠行走(SWAQ)框架,一个单阶段端到端方法,利用下一步的特权物理重建,在策略学习过程中塑造循环历史表示的保留内容,而部署时仅使用直接的历史到动作路径。在对齐训练设置下,SWAQ比最强的非本体感觉基线DWAQ高出15.0%的峰值平均地形等级,且每控制步推理乘加操作减少44.4%。层间探测显示,重构的物理变量相关信息在线性可解性上可一直保持到动作输出前一层。补充理论分析将特权变量可恢复性与基于历史和特权信息策略类间的可实现回报差距联系起来。结果表明,语义目标可塑造学习过程,无需对应部署控制器的结构分解。

原文摘要 · Abstract (English)

Partially observable locomotion requires a policy to act when task-relevant properties of the robot--environment state are not fully specified by instantaneous observations. Existing approaches often address this challenge by explicitly estimating missing physical variables or processing extended observation histories through structured architectures. We take a different view: partial observability is fundamentally an information-retention problem. The decisive question is not how task-relevant information enters the network, but whether the policy's internal state retains it. Guided by this perspective, we propose SleepWalking for Robot Locomotion (SWAQ), a one-stage end-to-end framework that uses next-step privileged physical reconstruction to shape what a recurrent history representation retains during policy learning, while the deployed actor uses only a direct history-to-action pathway. Under aligned training settings, SWAQ achieves a 15.0\% higher peak mean terrain level than DWAQ, the strongest non-exteroceptive baseline, while using 44.4\% fewer inference MACs per control step. Layerwise probes further show that information associated with the reconstructed physical variables remains linearly decodable through the policy head up to the layer preceding the action output. Complementary theoretical analysis relates privileged-variable recoverability to the achievable-return gap between history-based and privileged-information policy classes. These results suggest that semantic objectives can structure learning without requiring a corresponding architectural decomposition of the deployed controller.

机器人行走强化学习状态保持端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。