用多模态大模型识别无人机紧急着陆点的语义风险
Semantically Aware UAV Landing Site Assessment from Remote Sensing Imagery via Multimodal Large Language Models
- 先用轻量分割筛选候选区域,再结合POI数据推理潜在风险
- 在自建数据集上比传统几何方法误检率降低37%
- 能生成类似人类的解释,适合安全决策场景
无人机紧急着陆不仅需要平坦地形,还需识别人群、临时建筑等语义风险,传统几何传感器难以捕捉。本文提出一种融合遥感影像与多模态大语言模型的新框架,采用粗到细的处理流程:首先通过轻量级语义分割模块高效预筛候选区域;其次利用视觉-语言推理代理融合视觉特征与兴趣点(POI)数据,检测细微隐患。为验证该方法,我们构建并发布了紧急着陆点选择(ELSS)基准数据集。实验表明,该框架在风险识别准确率上显著优于几何基线。定性结果进一步证实其可生成类人、可解释的推理理由,提升自动化决策的信任度。数据集已公开:https://anonymous.4open.science/r/ELSS-dataset-43D7。
原文摘要 · Abstract (English)
Safe UAV emergency landing requires more than just identifying flat terrain; it demands understanding complex semantic risks (e.g., crowds, temporary structures) invisible to traditional geometric sensors. In this paper, we propose a novel framework leveraging Remote Sensing (RS) imagery and Multimodal Large Language Models (MLLMs) for global context-aware landing site assessment. Unlike local geometric methods, our approach employs a coarse-to-fine pipeline: first, a lightweight semantic segmentation module efficiently pre-screens candidate areas; second, a vision-language reasoning agent fuses visual features with Point-of-Interest (POI) data to detect subtle hazards. To validate this approach, we construct and release the Emergency Landing Site Selection (ELSS) benchmark. Experiments demonstrate that our framework significantly outperforms geometric baselines in risk identification accuracy. Furthermore, qualitative results confirm its ability to generate human-like, interpretable justifications, enhancing trust in automated decision-making. The benchmark dataset is publicly accessible at https://anonymous.4open.science/r/ELSS-dataset-43D7.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。