让无人机通过语言指令自动定位并放置物体,成功率超87%。
AERMANI-PLACE: Language Guided Object Placement with Aerial Manipulators

- 用图像编辑模型生成视觉标记,指示语言指令中的放置位置
- 在100个任务中平均成功率达87%,真实场景下仍保持72%成功率
- 适合需要自然语言交互的无人机智能操作场景
物体放置是空中操作任务的基础,但现有系统通常需用户以度量坐标明确指定位置,这不直观且需理解坐标系与场景几何,难以实际应用。相比之下,人类常通过语言与指向动作结合表达空间目标。受此启发,本文提出AERMANI-PLACE框架,实现基于语言的空中机械臂物体放置。给定场景图像与自然语言指令,图像编辑模型生成含视觉标记的修改后场景图,该标记通过深度观测映射到物理环境,恢复出度量级放置点,并由空中机械臂生成并执行放置轨迹。在100个语言引导放置任务上评估,方法在测试集上平均成功率达87%,在真实空中操作平台上的平均成功率为72%。视频演示:https://youtu.be/SgwwgLBsv0g
原文摘要 · Abstract (English)
Object placement is a fundamental component of aerial manipulation tasks, yet existing systems typically require the desired placement position to be specified explicitly in metric coordinates. Such interfaces are not intuitive and require users to reason about coordinate frames and scene geometry, making them difficult to use in practical deployments. In contrast, humans often communicate spatial goals through a combination of language and pointing gestures. Inspired by this observation, we present AERMANI-PLACE, a framework for language-guided object placement with aerial manipulators. Given a scene image and a natural language instruction, an image editing model generates a modified version of the scene containing a visual marker that indicates where the object should be placed. This marker is then grounded into the physical environment using depth observations to recover a metric place point, after which a placement trajectory is generated and executed by the aerial manipulator. We evaluate the proposed approach on a test set of 100 language-guided placement tasks and demonstrate successful execution on a real aerial manipulation platform. Experimental results show that the proposed method reliably infers placement locations from language instructions with an average success rate of 87\% on the test-set and transfers effectively to real-world aerial manipulation with an average success rate of 72\%. Video: https://youtu.be/SgwwgLBsv0g
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。