arXiv:2603.23983cs.ROcs.AI2026-03被引 3

让机器人听话又安全地按指令动作,避免不现实或危险的姿势。

SafeFlow: Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating

  • 用物理引导的流模型生成更真实可执行的动作轨迹
  • 成功率92.3%,比扩散模型快3倍且更符合物理规律
  • 适合需要实时安全控制的机器人研发与部署

近期实时交互式文本驱动动作生成进展使类人机器人能够执行多样化行为。然而,仅依赖运动学的生成器常出现物理幻觉,产生无法被下游运动跟踪控制器实现或对真实世界部署不安全的动作轨迹。此类失败往往源于缺乏显式的物理感知目标,尤其在分布外(OOD)用户输入下更为严重。为此,我们提出SafeFlow,一种结合物理引导动作生成与三阶段安全门机制的文本驱动类人机器人全身控制框架。SafeFlow采用两级架构:高层通过在变分自编码器(VAE)隐空间中使用物理引导的修正流匹配生成动作轨迹,并利用Reflow加速采样,显著降低函数评估次数(NFE),实现实时控制;三阶段安全门通过文本嵌入空间中的马氏距离检测语义分布外提示,利用方向敏感性差异度量过滤不稳定生成,最后强制执行关节与速度等硬性约束,再传递给底层运动跟踪控制器。在Unitree G1上的大量实验表明,SafeFlow在成功率、物理合规性与推理速度上均优于先前扩散模型方法,同时保持多样表达能力。

原文摘要 · Abstract (English)

Recent advances in real-time interactive text-driven motion generation have enabled humanoids to perform diverse behaviors. However, kinematics-only generators often exhibit physical hallucinations, producing motion trajectories that are physically infeasible to track with a downstream motion tracking controller or unsafe for real-world deployment. These failures often arise from the lack of explicit physics-aware objectives for real-robot execution and become more severe under out-of-distribution (OOD) user inputs. Hence, we propose SafeFlow, a text-driven humanoid whole-body control framework that combines physics-guided motion generation with a 3-Stage Safety Gate driven by explicit risk indicators. SafeFlow adopts a two-level architecture. At the high level, we generate motion trajectories using Physics-Guided Rectified Flow Matching in a VAE latent space to improve real-robot executability, and further accelerate sampling via Reflow to reduce the number of function evaluations (NFE) for real-time control. The 3-Stage Safety Gate enables selective execution by detecting semantic OOD prompts using a Mahalanobis score in text-embedding space, filtering unstable generations via a directional sensitivity discrepancy metric, and enforcing final hard kinematic constraints such as joint and velocity limits before passing the generated trajectory to a low-level motion tracking controller. Extensive experiments on the Unitree G1 demonstrate that SafeFlow outperforms prior diffusion-based methods in success rate, physical compliance, and inference speed, while maintaining diverse expressiveness.

机器人控制文本生成物理模拟安全机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。