arXiv:2603.01324cs.CV2026-03

对比了开放词汇与监督学习在灾后图像理解中的表现,发现监督方法更可靠。

Open-Vocabulary vs Supervised Learning Methods for Post-Disaster Visual Scene Understanding

  • 用开放词汇模型减少对标注数据的依赖,利用视觉语言预训练
  • 监督学习在小物体和边界分割上表现更好,尤其在杂乱场景中
  • 适合需要高精度、有标注数据的灾后应急响应场景

航拍图像对大规模灾后损毁评估至关重要。由于场景杂乱、视觉差异大及跨事件域偏移严重,自动化解读仍具挑战性;而传统监督方法依赖昂贵且任务特定的标注,覆盖范围有限。近期开放词汇与基础视觉模型提供了新路径,通过减少对固定标签集和大量标注的依赖,利用大规模预训练与视觉语言表征。这些特性在灾后场景中尤为适用,因视觉概念模糊且数据稀缺。本文对比了监督学习与开放词汇视觉模型在灾后场景理解中的表现,涵盖语义分割与目标检测,使用FloodNet+、RescueNet、DFire和LADD等多个数据集。分析性能趋势、失败模式与实际权衡,揭示其在真实灾害响应中的适用性。所有评估基准中最显著的发现是:当标签空间固定且有标注数据时,监督训练仍是更可靠的方案,尤其在小物体识别与复杂场景边界精确划分方面表现更优。

原文摘要 · Abstract (English)

Aerial imagery is critical for large-scale post-disaster damage assessment. Automated interpretation remains challenging due to clutter, visual variability, and strong cross-event domain shift, while supervised approaches still rely on costly, task-specific annotations with limited coverage across disaster types and regions. Recent open-vocabulary and foundation vision models offer an appealing alternative, by reducing dependence on fixed label sets and extensive task-specific annotations. Instead, they leverage large-scale pretraining and vision-language representations. These properties are particularly relevant for post-disaster domains, where visual concepts are ambiguous and data availability is constrained. In this work, we present a comparative evaluation of supervised learning and open-vocabulary vision models for post-disaster scene understanding, focusing on semantic segmentation and object detection across multiple datasets, including FloodNet+, RescueNet, DFire, and LADD. We examine performance trends, failure modes, and practical trade-offs between different learning paradigms, providing insight into their applicability for real-world disaster response. The most notable remark across all evaluated benchmarks is that supervised training remains the most reliable approach (i.e., when the label space is fixed and annotations are available), especially for small objects and fine boundary delineation in cluttered scenes.

灾后评估语义分割开放词汇监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。