用语义引导的视角规划,让无人机更快在复杂环境中找到目标。
STEM: Semantic Target Search and Exploration using MAVs in Cluttered Environments

- 设计组合式视角规划器,优先选择可能发现目标的观测位置。
- 通过语义信息增益计算,使探索路径更高效,搜索时间显著缩短。
- 支持真实场景约束,适合救援、搜救等实际应用。
自主目标搜索对微小型飞行器(MAVs)在应急响应和救援任务中的部署至关重要。现有方法要么侧重于结构化环境中的2D语义导航——在复杂3D场景中效果有限;要么关注杂乱空间的机器人探索——但缺乏高效目标搜索所需的语义推理能力。本文提出一种新框架,利用语义引导的视角规划器,在非结构化3D环境中最小化目标搜索与探索时间。我们开发了一种组合式规划器,通过优先选择可能导向目标的视角生成高效探索计划。为引导规划器向目标靠近,构建了主动感知流水线,将已观测物体的语义优先级传播至邻近前缘体素,以计算前缘视角的语义信息增益。此外,演示了如何使用基于大语言模型(LLM)的相似性得分作为语义优先级输入。在两个不同仿真环境中的评估显示,该方法始终优于基线,快速定位目标且探索时间合理。真实世界实验进一步验证了该方法在电池寿命有限、传感器范围小、语义不确定性等实际约束下的有效性。
原文摘要 · Abstract (English)
Autonomous target search is crucial for deploying Micro Aerial Vehicles (MAVs) in emergency response and rescue missions. Existing approaches either focus on 2D semantic navigation in structured environments -- which is less effective in complex 3D settings, or on robotic exploration in cluttered spaces -- which often lacks the semantic reasoning needed for efficient target search. This paper overcomes these limitations by proposing a novel framework that utilizes a semantically-guided viewpoint planner to minimize target search and exploration time in unstructured 3D environments using an MAV. Specifically, we develop a combinatorial planner that generates efficient semantic exploration plans by prioritizing viewpoints that likely lead to the target. To guide the planner towards the target, an active perception pipeline is developed that propagates semantic priorities of observed objects into neighboring frontier voxels for computing semantic information gains of frontier viewpoints. In addition, we demonstrate how LLM-based similarity scores can be leveraged as semantic priority input to our pipeline. Evaluations in two distinct simulation environments show that the proposed method consistently outperforms baselines by quickly finding the target while maintaining reasonable exploration times. Real-world experiments with an MAV further demonstrate the method's ability to handle practical constraints like limited battery life, small sensor range, and semantic uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。