arXiv:2603.25038cs.RO2026-03被引 3

将地面机械臂模型迁移到飞行机器人,靠物理引导和合成数据提升成功率。

$π$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial Manipulation

  • 用物理约束引导策略采样,解决飞行控制与地面模型不匹配问题。
  • 合成导航数据使飞行任务成功率从81%提升至100%,真实抓取成功率达50%。
  • 适合做无人机抓取、复杂任务组合的科研人员或工程团队参考。

视觉-语言-动作(VLA)模型如π₀在多种固定基机械臂上表现出优异泛化能力,但将其迁移到飞行平台仍面临挑战,因固定基臂的准静态动力学与飞行器的欠驱动、高度动态特性存在根本差异。本文提出AirVLA系统,研究预训练VLA在空中抓放任务中的可迁移性。发现视觉表征可有效迁移,但飞行所需控制动力学难以直接继承。为弥合这一“动力学鸿沟”,提出负载感知引导机制,将负载约束直接注入策略的流匹配采样过程。针对数据稀缺问题,引入高斯点云渲染流水线生成导航训练数据。通过累计460次真实实验验证:该合成数据是性能关键,使导航任务成功率从直接微调遥操作数据的81%提升至100%;推理时干预的负载感知引导将抓放任务成功率从23%提升至50%。在长时序组合任务中,整体成功率达62%。结果表明,结合适当数据增强与物理引导,预训练的操控型VLA可成功迁移至空中操控与导航,并支持任务组合。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models such as $π_0$ have demonstrated remarkable generalization across diverse fixed-base manipulators. However, transferring these foundation models to aerial platforms remains an open challenge due to the fundamental mismatch between the quasi-static dynamics of fixed-base arms and the underactuated, highly dynamic nature of flight. In this work, we introduce AirVLA, a system that investigates the transferability of manipulation-pretrained VLAs to aerial pick-and-place tasks. We find that while visual representations transfer effectively, the specific control dynamics required for flight do not. To bridge this "dynamics gap" without retraining the foundation model, we introduce a Payload-Aware Guidance mechanism that injects payload constraints directly into the policy's flow-matching sampling process. To overcome data scarcity, we further utilize a Gaussian Splatting pipeline to synthesize navigation training data. We evaluate our method through a cumulative 460 real-world experiments which demonstrate that this synthetic data is a key enabler of performance, unlocking 100% success in navigation tasks where directly fine-tuning on teleoperation data alone attains 81% success. Our inference-time intervention, Payload-Aware Guidance, increases real-world pick-and-place task success from 23% to 50%. Finally, we evaluate the model on a long-horizon compositional task, achieving a 62% overall success rate. These results suggest that pre-trained manipulation VLAs, with appropriate data augmentation and physics-informed guidance, can transfer to aerial manipulation and navigation, as well as the composition of these tasks.

飞行操控模型迁移物理引导合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。