arXiv:2607.22997cs.ROcs.AI2026-07

用AMD生态实现端到端物理智能操控,打破英伟达垄断

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

论文配图:Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
图 1 · 摘自论文原文
  • 全栈AMD加速,ROCm+PyTorch打通训练到部署
  • 实测在真实机械臂上运行语言控制抓取任务
  • 支持3D高斯泼溅+物理引擎生成仿真数据

物理人工智能——将大视觉-语言-动作(VLA)模型与具身智能体结合,在真实世界中执行任务——已成为AI新前沿。本文提出一个端到端、全AMD加速的技术栈,涵盖数据中心训练芯片、Radeon PRO仿真渲染显卡和Ryzen AI边缘计算,统一于开源ROCm软件平台。演示了四项技术:(1) 使用SmolVLA训练的模拟到真实操控流程,部署于真实Franka机械臂;(2) 基于语义的语言引导物体选择任务(“三选一”);(3) 真实场景3D高斯泼溅重建与Genesis物理引擎融合的实时到模拟合成数据生成;(4) 多硬件平台上的四足与人形机器人运动强化学习大规模基准测试。所有流程均原生运行于ROCm + PyTorch环境,在RDNA4(Radeon AI PRO R9700)和RDNA3.5(Radeon PRO W7900)硬件上可复现,并可在免费Radeon Cloud平台部署。

原文摘要 · Abstract (English)

Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged as the next major frontier for AI, echoed by industry leaders such as Jensen Huang (``the next big thing is Physical AI, AI with a body,'' GTC Paris, June 2025) and Dr. Lisa Su (`we're entering the world of Physical AI ... this is where AI enters the real world,' CES 2026). This paper presents an end-to-end, fully AMD-accelerated technology stack for embodied manipulation, spanning data-center training silicon, Radeon PRO simulation/rendering GPUs, and Ryzen AI edge compute, unified by the open ROCm software stack. We demonstrate that training and deploying VLA-based manipulation policies does not require a CUDA-locked ecosystem. Four progressive demonstrations are presented: (1) a Sim-to-Real manipulation pipeline trained with SmolVLA and deployed on a physical Franka arm; (2) a semantic, language-grounded object-selection task (`one-of-three'); (3) a Real2Sim synthetic-data generation pipeline that fuses 3D Gaussian Splatting (3DGS) reconstructions of real scenes with the Genesis physics engine; and (4) large-scale reinforcement learning for quadruped and humanoid locomotion benchmarked across multiple hardware platforms. All pipelines run natively on ROCm + PyTorch on RDNA4 (Radeon AI PRO R9700) and RDNA3.5 (Radeon PRO W7900) hardware and are reproducible on the free Radeon Cloud Platform.

物理AIAMD生态具身智能仿真生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。