让无人机能预判自身动作后果,实现三维空间自由想象。
AirScape: An Aerial Generative World Model with Motion Controllability
- 构建首个六自由度飞行世界模型,基于视觉与动作意图预测未来画面。
- 在11000组视频-意图数据上训练,运动对齐指标提升超50%。
- 适合研究智能体空间推理、无人机自主导航的开发者与学者。
如何让智能体在三维空间中预测自身运动意图的结果,是具身智能的核心挑战。为探索通用的空间想象能力,我们提出AirScape,首个专为六自由度飞行代理设计的世界模型。AirScape基于当前视觉输入和运动意图,预测未来的观测序列。我们构建了一个用于训练和测试的航空世界模型数据集,包含11,000个视频-意图配对,涵盖多样化的无人机动作与广泛场景的第一人称视角视频,标注过程耗时超过1,000小时。通过两阶段训练流程,将一个初始无具身空间知识的基础模型,转化为可由运动意图控制且符合物理时空约束的世界模型。实验表明,AirScape在三维空间想象能力上显著优于现有基础模型,尤其在反映运动对齐的指标上提升超过50%。项目地址:https://embodiedcity.github.io/AirScape/。
原文摘要 · Abstract (English)
How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general spatial imagination capability, we present AirScape, the first world model designed for six-degree-of-freedom aerial agents. AirScape predicts future observation sequences based on current visual inputs and motion intentions. Specifically, we construct a dataset for aerial world model training and testing, which consists of 11k video-intention pairs. This dataset includes first-person-view videos capturing diverse drone actions across a wide range of scenarios, with over 1,000 hours spent annotating the corresponding motion intentions. Then we develop a two-phase schedule to train a foundation model--initially devoid of embodied spatial knowledge--into a world model that is controllable by motion intentions and adheres to physical spatio-temporal constraints. Experimental results demonstrate that AirScape significantly outperforms existing foundation models in 3D spatial imagination capabilities, especially with over a 50% improvement in metrics reflecting motion alignment. The project is available at: https://embodiedcity.github.io/AirScape/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。