首个面向空中机械臂的视觉-语言-动作基准,解决飞行操控难题
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
- 构建基于物理的仿真环境与3000条手动操作数据集
- 验证现有模型在飞行平台上的可行性并揭示其能力边界
- 适合研究通用空中机器人与多模态智能的学者使用
尽管视觉-语言-动作(VLA)模型在地面机器人中取得显著进展,但其在空中机械臂系统(AMS)中的应用仍属空白。由于AMS具有浮游基座动力学、无人机与机械臂强耦合以及任务多步长周期等特性,现有为静态或二维移动平台设计的VLA范式难以适用。为此,我们提出首个专为航空操纵设计的VLA基准——AIR-VLA。构建了基于物理的仿真环境,并发布高质量多模态数据集,包含3000条人工遥操作示范,覆盖基座操控、物体与空间理解、语义推理及长时程规划。基于该平台,系统评估主流VLA模型与先进VLM模型。实验不仅验证了将VLA范式迁移至空中系统的可行性,更通过针对空中任务定制的多维度指标,揭示当前模型在无人机机动性、机械臂控制和高层规划方面的能力与局限。AIR-VLA为通用空中机器人研究建立了标准化测试平台与数据基础。资源将在 https://github.com/SpencerSon2001/AIR-VLA 公开。
原文摘要 · Abstract (English)
While Vision-Language-Action (VLA) models have achieved remarkable success in ground-based embodied intelligence, their application to Aerial Manipulation Systems (AMS) remains a largely unexplored frontier. The inherent characteristics of AMS, including floating-base dynamics, strong coupling between the UAV and the manipulator, and the multi-step, long-horizon nature of operational tasks, pose severe challenges to existing VLA paradigms designed for static or 2D mobile bases. To bridge this gap, we propose \textbf{AIR-VLA}, the first VLA benchmark specifically tailored for aerial manipulation. We construct a physics-based simulation environment and release a high-quality multimodal dataset comprising 3000 manually teleoperated demonstrations, covering base manipulation, object \& spatial understanding, semantic reasoning, and long-horizon planning. Leveraging this platform, we systematically evaluate mainstream VLA models and state-of-the-art VLM models. Our experiments not only validate the feasibility of transferring VLA paradigms to aerial systems but also, through multi-dimensional metrics tailored to aerial tasks, reveal the capabilities and boundaries of current models regarding UAV mobility, manipulator control, and high-level planning. \textbf{AIR-VLA} establishes a standardized testbed and data foundation for future research in general-purpose aerial robotics. The resource of AIR-VLA will be available at https://github.com/SpencerSon2001/AIR-VLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。