arXiv:2605.19728cs.CV2026-05

用惯性控制信号生成可控的无人机视频,让虚拟飞行训练更真实高效。

Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls

论文配图:Aero-World: Action-Conditioned Aerial Video Generation from Inertial Controls
图 1 · 摘自论文原文
  • 将惯性动作信号注入预训练模型,实现对无人机视频的精细控制。
  • 在基准测试中动作对齐度提升至63.6,运动稳定性显著增强。
  • 适合需要高精度物理一致性的无人机算法训练与评估场景。

基础视频模型虽视觉效果出色,但在具身智能中的应用受限,因其主要基于自然语言训练,而非底层控制信号。这一限制在空中飞行中尤为明显,因6自由度运动难以控制,微小姿态误差会导致轨迹漂移。通过生成遵循细粒度惯性动作的空中视频,可为无人机代理提供可控制的仿真替代数据,支持规模化训练与评估。为此,我们提出Aero-World,一种将预训练图像到视频扩散模型转化为可控空中视频生成器的方法。该方法通过动作令牌流将平移加速度和角速度序列注入预训练潜在扩散变压器。一个独立训练的冻结潜空间物理探测器,在LoRA微调过程中提供可微分的惯性一致性监督,避免了昂贵的视频解码。我们进一步提出AeroBench基准,用于评估生成无人机视频是否符合低层动作信号。AeroBench使用动作对齐分数(AAS)衡量与指令惯性动作的一致性,以及物理一致性率(PCR)衡量时序运动稳定性。在AeroBench上,Aero-World将平均AAS从57.7提升至63.6,相比AirScape表现出更优的质量-控制权衡:FVD更低(596.5 vs. 1058.6),SSIM更高(0.595 vs. 0.505),Flow-IMU相关性更强(0.44 vs. 0.20)。结果表明,冻结物理探测器监督是实现更动作对齐空中运动的有效机制。

原文摘要 · Abstract (English)

Foundation video models produce visually impressive results, but their use in embodied AI remains limited because they are primarily trained on natural language rather than low-level control signals. This limitation is especially pronounced for aerial flight, where motion occurs in unconstrained 6-DoF space and small errors in ego-motion can produce large trajectory drift. Generating aerial videos that follow fine-grained inertial actions can support scalable training and evaluation of aerial agents by providing a controllable proxy for real-world or expensive simulation data. To address this problem, we propose \textbf{Aero-World}, a method for converting a pretrained image-to-video diffusion model into a controllable aerial video generator. Aero-World injects sequences of translational acceleration and angular velocity into a pretrained latent diffusion transformer through an action-token stream. A frozen latent-space Physics Probe, trained independently on real video--IMU pairs, provides differentiable inertial-consistency supervision during LoRA finetuning while avoiding computationally expensive video decoding. We further propose \textbf{AeroBench}, a benchmark for evaluating whether generated drone videos adhere to low-level action signals. AeroBench uses Action Alignment Score (AAS) to measure agreement with commanded inertial actions and Physical Consistency Rate (PCR) to measure temporal motion stability. On AeroBench, Aero-World improves mean AAS from 57.7 to 63.6 over action-only finetuning and gives a stronger quality-control trade-off than AirScape, with lower FVD (596.5 vs. 1058.6), higher SSIM (0.595 vs. 0.505), and higher Flow-IMU correlation (0.44 vs. 0.20). These results suggest that frozen Physics Probe supervision is a practical mechanism for adapting pretrained video generators toward more action-aligned aerial motion.

视频生成无人机物理对齐扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。