arXiv:2511.01472cs.RO2025-11被引 7

用结构化提示让视觉语言模型安全操控无人机抓取物体

AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models

  • 将语言指令与安全约束转为结构化提示,分步推理
  • 通过预定义飞行安全技能库,避免幻觉和不安全动作
  • 无需微调,可泛化到新任务、新物体和新环境

视觉-语言模型(VLM)的快速发展为机器人控制带来了新可能,自然语言可表达操作目标,视觉反馈则连接感知与动作。然而,直接在空中机械臂上部署VLM驱动策略仍存在安全隐患且不可靠,因生成动作常不一致、易幻觉,且动态上不可行。本文提出AERMANI-VLM,首个无需任务微调即可适配预训练VLM用于空中操作的框架。该框架将自然语言指令、任务上下文与安全约束编码为结构化提示,引导模型生成自然语言的分步推理过程。该推理结果用于从预定义的离散飞行安全技能库中选择动作,确保执行可解释且时间上一致。通过解耦符号推理与物理动作,AERMANI-VLM有效缓解幻觉指令,防止不安全行为,实现鲁棒任务完成。我们在仿真与真实硬件上验证了该框架在多步骤拾取放置任务中的表现,证明其对未见过的指令、物体和环境具有强泛化能力。

原文摘要 · Abstract (English)

The rapid progress of vision--language models (VLMs) has sparked growing interest in robotic control, where natural language can express the operation goals while visual feedback links perception to action. However, directly deploying VLM-driven policies on aerial manipulators remains unsafe and unreliable since the generated actions are often inconsistent, hallucination-prone, and dynamically infeasible for flight. In this work, we present AERMANI-VLM, the first framework to adapt pretrained VLMs for aerial manipulation by separating high-level reasoning from low-level control, without any task-specific fine-tuning. Our framework encodes natural language instructions, task context, and safety constraints into a structured prompt that guides the model to generate a step-by-step reasoning trace in natural language. This reasoning output is used to select from a predefined library of discrete, flight-safe skills, ensuring interpretable and temporally consistent execution. By decoupling symbolic reasoning from physical action, AERMANI-VLM mitigates hallucinated commands and prevents unsafe behavior, enabling robust task completion. We validate the framework in both simulation and hardware on diverse multi-step pick-and-place tasks, demonstrating strong generalization to previously unseen commands, objects, and environments.

无人机操作视觉语言模型结构化推理零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。