arXiv:2605.03269cs.ROcs.AI2026-05被引 16

RDLX-1让机器人在复杂操作中表现更优,尤其在高自由度人形机器人任务中成功率达86.8%。

RLDX-1 Technical Report

论文配图:RLDX-1 Technical Report
图 1 · 摘自论文原文
  • 采用多流动作变换器架构,融合视觉、语言与动作模态,实现跨模态统一建模。
  • 在人形机器人复杂任务中成功率高达86.8%,远超π₀.₅和GR00T N1.₆的约40%。
  • 专为灵巧操作设计,适合需要长期记忆与物理感知的现实世界应用。

尽管视觉-语言-动作模型(VLAs)已通过预训练视觉-语言模型继承的通用智能(如广泛场景理解与语言条件泛化)取得显著进展,但在需要更广功能能力(如运动感知、长期记忆、物理传感)的复杂现实任务中仍表现不足。为此,我们提出RLDX-1,一种基于多流动作变换器(MSAT)的通用灵巧操作机器人策略,该架构通过模态专用流与跨模态联合自注意力机制整合异构信息。RLDX-1进一步结合系统级设计,包括稀有操作场景的数据合成、针对人类式操作的学习流程优化,以及实时部署的推理加速。实证评估表明,RLDX-1在仿真基准与需广泛功能能力的真实任务中持续优于近期前沿VLAs(如π₀.₅和GR00T N1.₆)。特别是在ALLEX人形机器人任务中,成功率达86.8%,而π₀.₅与GR00T N1.₆约为40%,凸显其在多样功能需求下控制高自由度人形机器人的能力。这些结果使RLDX-1成为迈向可靠、复杂、接触丰富且动态现实灵巧操作的有前景一步。

原文摘要 · Abstract (English)

While Vision-Language-Action models (VLAs) have shown remarkable progress toward human-like generalist robotic policies through the versatile intelligence (i.e. broad scene understanding and language-conditioned generalization) inherited from pre-trained Vision-Language Models, they still struggle with complex real-world tasks requiring broader functional capabilities (e.g. motion awareness, long-term memory, and physical sensing). To address this, we introduce RLDX-1, a general-purpose robotic policy for dexterous manipulation built on the Multi-Stream Action Transformer (MSAT), an architecture that unifies these capabilities by integrating heterogeneous modalities through modality-specific streams with cross-modal joint self-attention. RLDX-1 further combines this architecture with system-level design choices, including data synthesis for rare manipulation scenarios, learning procedures specialized for human-like manipulation, and inference optimizations for real-time deployment. Through empirical evaluation, we show that RLDX-1 consistently outperforms recent frontier VLAs (e.g. $π_{0.5}$ and GR00T N1.6) across both simulation benchmarks and real-world tasks that require broad functional capabilities beyond general versatility. In particular, RLDX-1 shows superiority in ALLEX humanoid tasks by achieving success rates of 86.8% while $π_{0.5}$ and GR00T N1.6 achieve around 40%, highlighting the ability of RLDX-1 to control a high-DoF humanoid robot under diverse functional demands. Together, these results position RLDX-1 as a promising step toward reliable VLAs for complex, contact-rich, and dynamic real-world dexterous manipulation.

机器人灵巧操作多模态强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。