构建了超30小时高动态无人机视频数据集,助力世界模型精准模拟复杂3D环境。
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models

- 采集4.5M帧4K无人机视频,含精确6-DoF轨迹与自然语言描述
- 引入多阶段自动化流程,实现语义与几何对齐的高质量标注
- 适用于需要高动态视角适应性的无人机自主导航研究
近期世界模型在模拟物理现实方面展现出强大能力,成为具身智能的重要基础。对于无人机代理而言,准确预测复杂三维动态对在非受限环境中自主导航与稳健决策至关重要。然而,在典型无人机视角的高动态相机运动下,现有世界模型常难以保持时空物理一致性。主要原因是当前训练数据存在分布偏差:多数数据集仅呈现受限的2.5D运动模式,如地面约束的自动驾驶场景或相对平滑的人类中心第一人称视频,缺乏真实的高动态六自由度(6-DoF)无人机运动先验。为填补这一空白,我们提出MotionScape,一个大规模真实世界无人机视角视频数据集,专为世界建模设计。该数据集包含超过30小时的4K无人机视频,总计超过450万帧。其特点在于语义与几何对齐的训练样本,将多样真实无人机视频紧密关联于精确6-DoF相机轨迹及细粒度自然语言描述。为构建此数据集,我们开发了一套自动化多阶段处理流水线,集成基于CLIP的相关性过滤、时间分割、鲁棒视觉SLAM轨迹恢复以及大语言模型驱动的语义标注。大量实验表明,引入此类语义与几何对齐的标注显著提升了现有世界模型在模拟复杂3D动态和应对大视角变化方面的能力,从而增强无人机代理在复杂环境中的决策与规划性能。数据集已公开于https://github.com/Thelegendzz/MotionScape。
原文摘要 · Abstract (English)
Recent advances in world models have demonstrated strong capabilities in simulating physical reality, making them an increasingly important foundation for embodied intelligence. For UAV agents in particular, accurate prediction of complex 3D dynamics is essential for autonomous navigation and robust decision-making in unconstrained environments. However, under the highly dynamic camera trajectories typical of UAV views, existing world models often struggle to maintain spatiotemporal physical consistency. A key reason lies in the distribution bias of current training data: most existing datasets exhibit restricted 2.5D motion patterns, such as ground-constrained autonomous driving scenes or relatively smooth human-centric egocentric videos, and therefore lack realistic high-dynamic 6-DoF UAV motion priors. To address this gap, we present MotionScape, a large-scale real-world UAV-view video dataset with highly dynamic motion for world modeling. MotionScape contains over 30 hours of 4K UAV-view videos, totaling more than 4.5M frames. This novel dataset features semantically and geometrically aligned training samples, where diverse real-world UAV videos are tightly coupled with accurate 6-DoF camera trajectories and fine-grained natural language descriptions. To build the dataset, we develop an automated multi-stage processing pipeline that integrates CLIP-based relevance filtering, temporal segmentation, robust visual SLAM for trajectory recovery, and large-language-model-driven semantic annotation. Extensive experiments show that incorporating such semantically and geometrically aligned annotations effectively improves the ability of existing world models to simulate complex 3D dynamics and handle large viewpoint shifts, thereby benefiting decision-making and planning for UAV agents in complex environments. The dataset is publicly available at https://github.com/Thelegendzz/MotionScape
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。