arXiv:2606.26201cs.RO2026-06被引 1

用接触流统一建模人形机器人移动操作,实现高成功率的长时序任务编排。

OmniContact: Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation

论文配图:OmniContact: Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation
图 1 · 摘自论文原文
  • 提出接触流表示,用关键轨迹和接触信号统一描述动作
  • 在搬箱和推叠箱任务中分别达98.7%和76.5%成功率
  • 支持语义分解与自动恢复,适合复杂场景下的机器人操作

学习长时序人形机器人移动操作面临双重挑战:既要稳健执行元技能,又要实现带自主恢复能力的闭环链式组合。现有方法受限于显式交互表示难于高层规划,或隐式技能嵌入缺乏可解释性。本文提出 extbf{OmniContact},一种以接触流(CF)为核心的分层框架,其为紧凑的键位轨迹与时间序列二值接触信号组合。底层策略 extbf{CF-Track}学习统一的动作技能库,高层模块 extbf{CF-Gen}启发式生成未来接触流序列。为此,我们构建了基于动捕的OmniContact数据集,涵盖人形机器人交互行为。实验表明,该方法在 extit{Carry Box}任务上成功率达98.7%, extit{Push-Stack Boxes}任务达76.5%,相比基线平均提升40.9%(元技能)与66.5%(链式组合)。此外,框架可自然融合视觉语言模型,实现如将散落箱子排列成心形等语义驱动操作。

原文摘要 · Abstract (English)

Learning long-horizon humanoid loco-manipulation poses a dual challenge: it requires not only the robust execution of meta-skills but also their seamless, closed-loop chaining equipped with autonomous recovery. Existing approaches remain limited: explicit humanoid-object interaction representations offer precision but are notoriously difficult for high-level planning, whereas implicit skill embeddings are compact but lack the interpretability required for reliable composition. We propose \ours, a hierarchical framework centered on \textbf{contact flow (CF)}, a compact representation consisting of key body trajectories and time-series binary contact signals. Leveraging this shared interface, our low-level policy \textbf{CF-Track} learns a unified library of loco-manipulation skills, while our high-level module \textbf{CF-Gen} heuristically synthesizes future contact-flow sequences. To support this setting, we additionally collect the OmniContact dataset, a MoCap-based HOI corpus for humanoid loco-manipulation (Appendix~\ref{sec:dataset}). Together, they enable robust execution, autonomous failure recovery, and flexible composition of meta-skills for long-horizon tasks. Experiments show that OmniContact achieves \(98.7\%\) success on \textit{Carry Box} and \(76.5\%\) on \textit{Push-Stack Boxes}, outperforming prior baselines by average margins of \(40.9\%\) in meta-skill and \(66.5\%\) in skill chaining. Besides, our framework naturally integrates with VLMs for semantic task decomposition, enabling complex, semantically grounded loco-manipulation behaviors, such as arranging scattered boxes into a heart shape.

人形机器人动作链接触建模长时序控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。