arXiv:2602.21723cs.RO2026-02被引 6

用距离场统一表征环境,让机器人无参考地完成复杂交互任务。

LessMimic: Long-Horizon Humanoid Interaction with Unified Distance Field Representations

  • 以距离场几何特征驱动单一全身策略,摆脱动作参考依赖。
  • 在0.4到1.6倍尺度物体上成功率80%~100%,40个任务连续执行仍有效。
  • 支持纯视觉部署,无需运动捕捉,适合真实场景通用机器人。

能长期自主与物理环境交互的类人机器人是具身智能的核心目标。现有方法依赖动作参考或任务特定奖励,导致策略与特定物体形状紧密耦合,难以实现多技能通用。本文提出LessMimic,利用距离场(DF)提供统一的交互表征:通过表面距离、梯度和速度分解等几何线索,构建单个全身策略,无需动作参考;交互潜变量由变分自编码器(VAE)编码,并在强化学习中通过对抗性交互先验(AIP)后训练。采用DAgger式蒸馏,将距离场潜变量与本体深度特征对齐,实现无需运动捕捉的纯视觉部署。单一策略在PickUp和SitStand任务中,面对0.4x至1.6x尺度物体,成功率保持80%~100%,基线模型则显著下降;在5个任务实例轨迹上达成62.1%成功率,可稳定执行最多40个连续组合任务。通过基于局部几何而非示范进行交互建模,LessMimic为类人机器人在非结构化环境中实现泛化、技能组合与容错提供了可扩展路径。

原文摘要 · Abstract (English)

Humanoid robots that autonomously interact with physical environments over extended horizons represent a central goal of embodied intelligence. Existing approaches rely on reference motions or task-specific rewards, tightly coupling policies to particular object geometries and precluding multi-skill generalization within a single framework. A unified interaction representation enabling reference-free inference, geometric generalization, and long-horizon skill composition within one policy remains an open challenge. Here we show that Distance Field (DF) provides such a representation: LessMimic conditions a single whole-body policy on DF-derived geometric cues--surface distances, gradients, and velocity decompositions--removing the need for motion references, with interaction latents encoded via a Variational Auto-Encoder (VAE) and post-trained using Adversarial Interaction Priors (AIP) under Reinforcement Learning (RL). Through DAgger-style distillation that aligns DF latents with egocentric depth features, LessMimic further transfers seamlessly to vision-only deployment without motion capture (MoCap) infrastructure. A single LessMimic policy achieves 80--100% success across object scales from 0.4x to 1.6x on PickUp and SitStand where baselines degrade sharply, attains 62.1% success on 5 task instances trajectories, and remains viable up to 40 sequentially composed tasks. By grounding interaction in local geometry rather than demonstrations, LessMimic offers a scalable path toward humanoid robots that generalize, compose skills, and recover from failures in unstructured environments.

类人机器人距离场技能组合视觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。