arXiv:2603.07744cs.RO2026-03被引 5

让无人机通过语言指令精准放物,无需预设位置。

AeroPlace-Flow: Language-Grounded Object Placement for Aerial Manipulators via Visual Foresight and Object Flow

  • 用图像编辑生成目标场景,再转为3D空间中的可执行动作
  • 在真实飞行中实现75%成功率,支持自然语言控制
  • 无需训练,适配复杂环境下的语言指令放物任务

精确物体放置在空中操作中仍缺乏研究,现有系统多依赖预设目标坐标,侧重抓取与控制。但在现实场景中,用户更习惯用语言表达目标。本文提出AeroPlace-Flow,一种无需训练的语言引导空中物体放置框架,融合视觉前瞻、显式3D几何推理与物体运动流。给定物体和场景的RGB-D观测及自然语言指令,该方法首先利用图像编辑模型合成任务完整的目标图像,再通过深度对齐与物体中心推理,将想象配置映射到度量3D空间,推断出避障的物体运动流,将抓取物运送到符合语言描述与接触一致的放置位姿。最终动作通过标准轨迹跟踪在空中机械臂上执行。AeroPlace-Flow无需预定义位姿或任务特定训练,即可生成可执行放置目标。我们在大量仿真与真实实验中验证该方法,在硬件上实现多样空中场景下语言条件放置的可靠表现,平均成功率达75%。

原文摘要 · Abstract (English)

Precise object placement remains underexplored in aerial manipulation, where most systems rely on predefined target coordinates and focus primarily on grasping and control. Specifying exact placement poses, however, is cumbersome in real-world settings, where users naturally communicate goals through language. In this work, we present AeroPlace-Flow, a training-free framework for language-grounded aerial object placement that unifies visual foresight with explicit 3D geometric reasoning and object flow. Given RGB-D observations of the object and the placement scene, along with a natural language instruction, AeroPlace-Flow first synthesizes a task-complete goal image using image editing models. The imagined configuration is then grounded into metric 3D space through depth alignment and object-centric reasoning, enabling the inference of a collision-aware object flow that transports the grasped object to a language and contact-consistent placement configuration. The resulting motion is executed via standard trajectory tracking for an aerial manipulator. AeroPlace-Flow produces executable placement targets without requiring predefined poses or task-specific training. We validate our approach through extensive simulation and real-world experiments, demonstrating reliable language-conditioned placement across diverse aerial scenarios with an average success rate of 75% on hardware.

无人机操作语言导航3D推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。