arXiv:2606.12859cs.RO2026-06

分离飞行与机械臂控制,提升无人机操作的稳定性和任务完成率。

AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

论文配图:AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots
图 1 · 摘自论文原文
  • 分层双解码器设计,让飞行与抓取独立演进,避免信息干扰。
  • 在标准基准上任务完成率提升80.2%,达48.0分,优于所有基线方法。
  • 适合需要高精度协同控制的空中机器人系统研究者参考。

空中操作系统的端到端控制长期受制于平台级无人机运动与末端执行器级机械臂操作之间表示耦合的问题,二者在动作尺度、动力学特性及控制目标上存在显著差异。本文提出AIR-VLA+,一种专为空中操作设计的流匹配动作生成架构,包含级联双解码器与非对称特征级混合专家(MoE)。构建级联式操作与运动解码器,使无人机在运动时单向感知机械臂意图,实现工作流程协调,同时隔离飞行运动信息反向传播对机械臂稳定性的影响。针对无人机运动高度依赖高层语义且负责任务状态转换的特点,为运动解码器设计输入特征增强模块,引入隐式视觉抓取投影器以感知夹爪与物体间的交互状态,并注入压缩全局语义特征。在无人机运动解码器中部署隐式MoE架构,使不同运动专家在训练中自发呈现对不同任务阶段的容量偏好。通过特征流形上的密集软融合计算,赋予无人机运动更强的任务阶段适应性。在标准化的AIR-VLA基准测试中,该方法整体平均得分48.0,全面超越所有基线模型,相较于单头π₀.₅策略,任务完成率提升80.2%,有效缓解复合机器人的异构协同控制冲突。

原文摘要 · Abstract (English)

Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VLA+, a flow matching action generation architecture specifically designed for aerial manipulation, featuring cascaded dual-action decoders and an asymmetric feature-level Mixture of Experts (MoE). We construct cascaded manipulation and movement decoders, allowing the UAV to unidirectionally observe the manipulator's intent during movement to achieve workflow coordination, while isolating the impact of UAV movement information backpropagation on arm manipulation stability. Addressing the characteristic that UAV movement is highly dependent on high-level semantics and responsible for task state transitions in aerial manipulation, we design an input feature enhancement module for the UAV movement decoder. This module introduces an implicit visual grasp projector to perceive the interaction state between the gripper and the object, and injects compressed global semantic features. Within the UAV movement decoder, we deploy an implicit MoE architecture, enabling different movement experts to spontaneously exhibit capacity inclinations for various task stages during training. Through dense soft blending computation on the feature manifold, the UAV movement is endowed with stronger task-stage adaptability. Experiments on the standardized AIR-VLA benchmark demonstrate that our method comprehensively surpasses all baselines with an overall average score of 48.0. The overall task completion score improves by 80.2\% compared to the single-head $π_{0.5}$ policy, effectively mitigating the heterogeneous coordinated control conflicts of composite robots.

空中操作双解码器混合专家无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。