arXiv:2606.01851cs.RO2026-06

构建通用动作表示空间,让不同人形机器人共享同一套可解释的动作语义体系。

PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments

论文配图:PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
图 1 · 摘自论文原文
  • 用相位-姿态分解建模动作周期性,通过傅里叶系数捕捉循环结构。
  • 在多个机器人上实现动作检索准确率提升,下游任务成功率显著提高。
  • 适合需要跨平台迁移的机器人控制研究者,尤其关注动作可解释性场景。

学习高质量的动作嵌入空间是实现可扩展机器人策略学习的基础,但现有方法将动作潜在变量视为任务相关的中间表示,而非第一类表征。这导致潜在变量无结构、依赖具体机体、与运动语义关联弱,限制了可解释性、可控性和跨机器人迁移能力。本文将动作嵌入空间本身作为首要设计目标,认为下游策略性能源于表征质量。利用运动固有的周期性,我们将运动分解为基于FFT参数化的相位流形,以及依赖非周期性构型细节的姿态分支。结合运动语义蒸馏,这种分解结构生成了一个跨机体、可解释且机体无关的动作流形。将多个不同人形机器人锚定至共享的人类预训练流形后,实现了跨平台统一的动作嵌入空间,在跨机体动作检索中表现优异,并在下游机器人任务中获得稳定提升。

原文摘要 · Abstract (English)

Learning a good action embedding space is fundamental to scalable robot policy learning, yet existing methods treat action latents as task-specific intermediates rather than first-class representations. The resulting latents are unstructured, embodiment-specific, and weakly tied to motion semantics, limiting interpretability, controllability, and transferability across robots. We position the action embedding space itself as a first-class design target, with downstream policy quality emerging from representation quality. Exploiting motion's intrinsic periodicity, we factorize it into a phase manifold that captures cyclic structure via FFT-parametric coefficients, together with a pose branch that conditions the manifold on non-periodic configuration detail. Combined with motion-semantic distillation, this factorized structure yields a cross-embodiment motion manifold that is interpretable and embodiment-agnostic by design. Anchoring multiple humanoid robots to a shared human-pretrained manifold then produces a unified action embedding space across diverse platforms, achieving strong cross-embodiment retrieval and consistent gains on downstream robot tasks.

动作表示机器人迁移周期性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。