arXiv:2512.24321cs.CVcs.RO2025-12被引 9

统一处理多模态指令,让人形机器人实时灵活执行动作

UniAct: Unified Motion Generation and Action Streaming for Humanoid Robots

  • 用共享离散码本统一语言、音乐等多模态输入
  • 实现500毫秒内响应,零样本追踪成功率提升19%
  • 适用于真实场景下复杂指令的通用人形机器人

人形机器人长期目标是能灵活响应多样多模态指令。尽管控制技术进步,但高层感知与全身执行之间的衔接仍是瓶颈。现有方法难以将语言、音乐、轨迹等异构指令转化为稳定实时的动作。本文提出UniAct,一个两阶段框架,结合微调的多模态大模型与因果流式处理管道,使机器人在子500毫秒延迟内执行多模态指令。通过使用FSQ的共享离散码本统一输入,确保跨模态对齐,并将运动限制在物理合理的流形上。该方法在20小时的人形动作基准UniMoCap上验证,实现零样本追踪不完美参考动作的成功率提升19%。结果标志着向可响应、通用型人形助手迈出关键一步。

原文摘要 · Abstract (English)

A long-standing objective in humanoid robotics is the realization of versatile agents capable of following diverse multimodal instructions with human-level flexibility. Despite advances in humanoid control, bridging high-level multimodal perception with whole-body execution remains a significant bottleneck. Existing methods often struggle to translate heterogeneous instructions -- such as language, music, and trajectories -- into stable, real-time actions. Here we show that UniAct, a two-stage framework integrating a fine-tuned MLLM with a causal streaming pipeline, enables humanoid robots to execute multimodal instructions with sub-500 ms latency. By unifying inputs through a shared discrete codebook via FSQ, UniAct ensures cross-modal alignment while constraining motions to a physically grounded manifold. This approach yields a 19% improvement in the success rate of zero-shot tracking of imperfect reference motions. We validate UniAct on UniMoCap, our 20-hour humanoid motion benchmark, demonstrating robust generalization across diverse real-world scenarios. Our results mark a critical step toward responsive, general-purpose humanoid assistants capable of seamless interaction through unified perception and control.

人形机器人多模态生成实时控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。