用世界模型驱动的赛车智能体,探索具身智能的认知与物理极限。
Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent

- 以真实赛车数据训练世界模型,捕捉极限状态下的交互演化。
- 在仿真中实现88.3%的交互成功率,逼近车辆动力学与认知能力边界。
- 适合研究高动态系统安全控制与极限性能优化的学者。
具身人工智能旨在构建通过持续与物理世界交互来感知、推理并行动的智能体。然而,多数系统仍局限于保守安全范围或中等交互强度,对其在极端条件下的能力边界了解不足。自主赛车提供了一个严苛的测试平台,融合高频定位与感知、对抗性交互、接近饱和的车辆动力学及严格安全约束。现有系统虽追求高速表现,却极少协同建模与优化认知与物理极限。本文展示,以世界模型为中心的自主赛车智能体为探索这些耦合极限迈出关键一步。该框架从接近极限的成功与失败数据中学习预测性世界模型,捕捉交互演化、自身动态与可行运动边界,将世界状态构建、未来感知推理与近极限控制闭环融合。训练数据来自真实车辆自主赛车,车载系统在256.3 km/h速度和峰值侧向加速度26.8 m/s²下保持稳定定位与感知。在全尺度仿真赛车中,训练良好的世界模型中心智能体在多种挑战性场景中达到88.3%的交互成功率。世界模型与策略的闭环优化进一步提升了对认知-物理极限的利用、故障恢复能力以及跨不同条件与未知赛道的泛化性能。结果表明,一种边界感知的方法论正逐步形成:世界模型帮助具身智能体表征、预测并持续精炼其能力边界,为更安全的实际部署提供支撑。
原文摘要 · Abstract (English)
Embodied artificial intelligence aims to develop agents that perceive, reason, and act through continuous interaction with the physical world. However, most embodied systems are still evaluated within conservative safety margins or moderate interaction regimes, leaving their capability boundaries under extreme conditions insufficiently understood. Autonomous racing provides a stringent testbed by combining high-frequency localization and perception, adversarial interaction, near-saturated vehicle dynamics, and strict safety constraints. Existing systems push high-speed performance but rarely model and refine cognitive and physical limits jointly. Here we show that a world-model-centric autonomous racing agent provides a concrete step toward exploring these coupled limits. The framework learns predictive world models from near-limit successes and failures to capture interaction evolution, ego dynamics, and feasible-motion boundaries, coupling world-state construction, future-aware reasoning, and near-limit control in a closed-loop refinement process. Training data were collected from real-vehicle autonomous racing, where the onboard system maintained robust localization and perception at speeds up to 256.3 km/h and peak lateral acceleration of 26.8 m/s$^2$. In full-scale simulated racing, the well trained world-model-centric agent achieves an 88.3% interaction success rate across various challenging simulated racing scenarios. Closed-loop refinement of the world model and policy further improved utilization of cognitive-physical limits, recovery from failure modes, and generalization across varying conditions and unseen circuits. These results suggest a boundary-aware methodology in which world models help embodied agents represent, predict, and continually refine their capability boundaries for safer real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。