让无人机在杂乱环境里可靠抓取物体,结合语言指令与主动探索。
AeroGrab: A Unified Framework for Aerial Grasping in Cluttered Environments
- 基于语言指令和主动探索获取多视角,生成6自由度抓取候选。
- 通过避障评估框架筛选最优抓取姿势,实测在复杂场景中稳定抓取。
- 适合需要语言控制、动态环境感知的无人机抓取应用开发者。
在杂乱环境中实现可靠的空中抓取仍具挑战,主要源于遮挡和碰撞风险。现有空中操作流程多依赖质心抓取,缺乏抓取姿态生成模型、主动探索与语言级任务指令之间的整合,导致无法形成完整端到端系统。本文提出一个集成化管道,用于复杂环境中的可靠空中抓取。给定场景与语言指令后,系统识别目标物体并主动探索以获取更好视图。探索过程中,抓取生成网络为每个视角预测多个6-DoF抓取候选,再通过避障可行性框架进行评估,最终选择最优抓取并使用标准轨迹生成与控制方法执行。真实世界杂乱场景实验表明,该方法能实现鲁棒且可靠的抓取,验证了主动感知与可行性感知抓取选择相结合在空中操作中的有效性。
原文摘要 · Abstract (English)
Reliable aerial grasping in cluttered environments remains challenging due to occlusions and collision risks. Existing aerial manipulation pipelines largely rely on centroid-based grasping and lack integration between the grasp pose generation models, active exploration, and language-level task specification, resulting in the absence of a complete end-to-end system. In this work, we present an integrated pipeline for reliable aerial grasping in cluttered environments. Given a scene and a language instruction, the system identifies the target object and actively explores it to gain better views of the object. During exploration, a grasp generation network predicts multiple 6-DoF grasp candidates for each view. Each candidate is evaluated using a collision-aware feasibility framework, and the overall best grasp is selected and executed using standard trajectory generation and control methods. Experiments in cluttered real-world scenarios demonstrate robust and reliable grasp execution, highlighting the effectiveness of combining active perception with feasibility-aware grasp selection for aerial manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。