构建无人机航拍建筑入口数据集,提升震后搜救实时识别精度。
DRespNeT: A UAV Dataset and YOLOv8-DRN Model for Aerial Instance Segmentation of Building Access Points for Post-Earthquake Search-and-Rescue Missions
- 基于1080p航拍视频构建细粒度实例分割数据集
- 自研YOLOv8-DRN模型达92.7% mAP50,推理速度27 FPS
- 适用于灾后搜救机器人与人工协同,支持实时决策
近年来计算机视觉与深度学习的进步显著提升了灾害响应能力,尤其在快速评估地震后的城市环境方面。及时识别可进入的出入口和结构障碍物对高效搜救行动至关重要。为此,我们提出了DRespNeT,一个专为震后建筑环境航空实例分割设计的高分辨率数据集。与依赖卫星图像或粗粒度语义标注的现有数据集不同,DRespNeT基于高清(1080p)航拍影像,涵盖2023年土耳其地震及其他受灾区域的真实画面,提供多边形级实例分割标注。数据集包含28类关键目标:包括受损建筑、门、窗、缝隙等入口,多重碎屑层,救援人员、车辆及民众可见性。其精细标注可区分可通行与受阻区域,有助于优化应急规划与响应效率。使用基于YOLO的实例分割模型(特别是YOLOv8-seg)进行性能评估,结果表明该方法显著提升实时态势感知与决策能力。我们优化的YOLOv8-DRN模型在RTX-4090 GPU上实现92.7% mAP50,推理速度达27 FPS,满足实时操作需求。该数据集与模型可支持搜救团队与机器人系统,为增强人机协作、优化应急响应、改善幸存者结局提供基础。
原文摘要 · Abstract (English)
Recent advancements in computer vision and deep learning have enhanced disaster-response capabilities, particularly in the rapid assessment of earthquake-affected urban environments. Timely identification of accessible entry points and structural obstacles is essential for effective search-and-rescue (SAR) operations. To address this need, we introduce DRespNeT, a high-resolution dataset specifically developed for aerial instance segmentation of post-earthquake structural environments. Unlike existing datasets, which rely heavily on satellite imagery or coarse semantic labeling, DRespNeT provides detailed polygon-level instance segmentation annotations derived from high-definition (1080p) aerial footage captured in disaster zones, including the 2023 Turkiye earthquake and other impacted regions. The dataset comprises 28 operationally critical classes, including structurally compromised buildings, access points such as doors, windows, and gaps, multiple debris levels, rescue personnel, vehicles, and civilian visibility. A distinctive feature of DRespNeT is its fine-grained annotation detail, enabling differentiation between accessible and obstructed areas, thereby improving operational planning and response efficiency. Performance evaluations using YOLO-based instance segmentation models, specifically YOLOv8-seg, demonstrate significant gains in real-time situational awareness and decision-making. Our optimized YOLOv8-DRN model achieves 92.7% mAP50 with an inference speed of 27 FPS on an RTX-4090 GPU for multi-target detection, meeting real-time operational requirements. The dataset and models support SAR teams and robotic systems, providing a foundation for enhancing human-robot collaboration, streamlining emergency response, and improving survivor outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。