arXiv:2608.19375cs.RO2026-08

用神经简化动力学模型,让机器人在低成本模拟中高效训练,还能直接迁移到高精度仿真。

Learning the Right Abstraction: Neural Reduced Dynamics for Complex Robot Control

  • 构建神经简化动力学框架,分离状态变量与输入,实现高效模拟
  • 单个策略在三种地形上均优于专用策略,零样本泛化能力出色
  • 速度比高保真模拟快四个数量级,适合大规模强化学习

高保真具身智能仿真器能真实评估复杂机器人系统,但计算成本过高,难以用于大规模强化学习。本文主张使用更快速但精度较低的仿真,尤其是基于数据驱动的神经动力学模型。关键在于学习「正确的抽象」:一种保留控制相关物理特性的简化状态,同时支持高吞吐量策略学习。我们提出神经简化动力学(NRD)框架,将模型演化的状态与可输入或解析恢复的变量分离,完全在冻结的神经模型中训练策略,并在高保真仿真器中验证。两个案例研究涵盖三项任务:在刚性、崎岖和可变形连续体表示模型(CRM)地形上的HMMWV轨迹跟踪;以及一辆履带车及其前装机械臂的目标到达任务。所有策略均可成功迁移回高保真环境。一个仅在地形条件化动力学模型中训练、未接收地形输入的策略,在三类地形上均优于单地形专用策略,包括零样本的崎岖地形。量化结果:车辆100/100完成目标,机械臂97/100完成目标,无接触或关节越界。NRD模型的模拟速度比其替代的高保真场景快约四个数量级,使迭代式策略学习成为可能,证明了神经简化动力学是连接高精度但昂贵物理仿真与可扩展机器人学习之间的桥梁。

原文摘要 · Abstract (English)

High-fidelity embodied AI simulators provide realistic evaluation of complex robotic systems, but their computational cost limits their direct use for large-scale reinforcement learning campaigns. We advocate the use of less accurate but more expeditious simulations, which might draw on data-driven, e.g., neural dynamics, models. This contribution argues that the practical value of a neural dynamics model for complex robot control lies in learning the \emph{right abstraction}: a reduced state that preserves the control-relevant physics of the high-fidelity system while enabling high-throughput policy learning. We develop a neural reduced dynamics (NRD) framework that separates the state the model propagates from what can be supplied as an input or recovered analytically, trains policies entirely inside the frozen learned model, and validates them back in the high-fidelity simulator. Two case studies instantiate it across three control tasks: terrain-aware HMMWV trajectory tracking on rigid, bumpy and deformable Continuum Representation Model (CRM) terrain; and goal reaching for a stock tracked vehicle and its front-mounted articulated arm. Every policy transfers back to the high-fidelity simulator. A single policy trained inside the terrain-conditioned dynamics model, and given no terrain input of its own, attains lower median and mean tracking error than both single-terrain specialists on all three terrains, including zero-shot bumpy terrain. Quantitatively, the tracked vehicle reaches 100 of 100 goals and the arm 97 of 100, with zero contacts or joint-limit violations. The NRD models advance roughly four orders of magnitude faster in simulated time than the high-fidelity simulator scenes they replace, making iterative on-policy learning practical and supporting neural reduced dynamics as a bridge between accurate but expensive physics simulation and scalable robot learning.

机器人控制神经动力学仿真加速强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。