用视觉语言建模让无人机按指令飞行,实时生成平滑控制指令。
AeroAct: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight

- 基于视觉历史与语言指令预测局部飞行动作序列
- 在仿真和真实飞行中均实现目标追踪与搜索性能提升
- 适合需要自然语言控制的无人机自主导航场景
语言控制无人机飞行需将语义目标具身化,预判自身运动带来的视觉变化,并输出在快速变化的第一视角下仍平滑可执行的控制参考。现有方法多使用离散动作、高层航点或瞬时速度指令,难以提供动作如何影响未来观测的充分监督。本文提出AeroAct,首个面向真实飞行的以动作为中心的世界-动作模型(WAM)。该模型利用预训练视频扩散Transformer,从第一人称视觉历史、本体感知和语言中预测局部轨迹-动作片段。训练时使用未来第一视角帧作为密集结果监督,部署时直接解码动作而不生成视频。为获取对齐的视觉、状态、语言与动态可行动作数据,构建了基于DiffAero的管道,结合Isaac Lab与3D高斯溅射渲染器。还设计低成本手持采集设备,融合摄像头与运动估计重建类飞行视角轨迹,并引入自引导机制提升重叠轨迹片段间的时序一致性。闭环仿真与真实实验表明,时间视觉上下文显著提升目标追踪与物体搜索性能,且基于WAM的策略可在真实四旋翼上执行。
原文摘要 · Abstract (English)
Language-conditioned quadrotor flight requires a policy to ground semantic goals, anticipate the visual consequences of ego-motion, and output control references that remain smooth and dynamically executable under rapidly changing first-person views. Existing aerial vision-language navigation and vision-language-action methods commonly use discrete actions, high-level waypoints, or instantaneous velocity commands, which provide limited supervision about how flight actions change future observations. We present AeroAct, an action-centered world-action model (WAM) for quadrotor navigation. To the best of our knowledge, AeroAct is the first WAM instantiated and demonstrated for real-world aerial flight. The model adapts a pretrained video diffusion Transformer to predict local trajectory-action chunks from egocentric visual history, proprioception, and language. Future first-person frames are used during training as dense consequence supervision, while deployment directly decodes actions without generating future video. To obtain aligned visual, state, language, and dynamically feasible action data, we build a DiffAero-based pipeline with complementary Isaac Lab and 3D Gaussian splatting renderers. We further introduce a low-cost handheld collection device that couples camera observations with motion estimates to recreate flight-like egocentric trajectories, and a self-guidance procedure that improves temporal consistency across overlapping trajectory chunks. Closed-loop simulation and real-world experiments show that temporal visual context improves target tracking and object-search performance, and that WAM-based policies can be executed on a physical quadrotor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。