arXiv:2504.13776cs.CVeess.IV2025-04被引 1

用视觉变压器提升卫星火情检测,效果媲美甚至超越传统卷积网络。

Fighting Fires from Space: Leveraging Vision Transformers for Enhanced Wildfire Detection and Characterization

  • 采用视觉变压器(ViT)处理陆地卫星影像,融合局部与全局上下文信息。
  • 最优ViT模型准确率比基准CNN高0.92%,但定制化CNN-UNet仍更优。
  • 适合关注遥感火灾监测、想对比深度学习架构的科研与应用人员。

野火因人为气候变化在全球多地变得愈发频繁、剧烈且持续时间更长。当前灾害检测与响应系统难以应对长期野火季。已有研究证明,基于卫星图像训练的卷积神经网络(CNN)可实现高精度自动火情检测,但其训练成本高且仅依赖局部图像上下文。近年来,视觉变压器(ViTs)因其高效训练和兼具局部与全局上下文建模能力而受到关注。本文在已发布的陆地卫星8号(LandSat-8)影像数据集上验证,ViT在野火检测中表现优于经过良好训练的专用CNN,其中一模型准确率领先0.92%。然而,我们自研的基于CNN的UNet在各项指标上均最优,显示出其在图像任务中的持续优势。总体而言,ViT在野火检测中表现与CNN相当,但经充分调优的CNN-UNet仍为最佳方案,其交并比(IoU)达93.58%,较基线提升4.58%。

原文摘要 · Abstract (English)

Wildfires are increasing in intensity, frequency, and duration across large parts of the world as a result of anthropogenic climate change. Modern hazard detection and response systems that deal with wildfires are under-equipped for sustained wildfire seasons. Recent work has proved automated wildfire detection using Convolutional Neural Networks (CNNs) trained on satellite imagery are capable of high-accuracy results. However, CNNs are computationally expensive to train and only incorporate local image context. Recently, Vision Transformers (ViTs) have gained popularity for their efficient training and their ability to include both local and global contextual information. In this work, we show that ViT can outperform well-trained and specialized CNNs to detect wildfires on a previously published dataset of LandSat-8 imagery. One of our ViTs outperforms the baseline CNN comparison by 0.92%. However, we find our own implementation of CNN-based UNet to perform best in every category, showing their sustained utility in image tasks. Overall, ViTs are comparably capable in detecting wildfires as CNNs, though well-tuned CNNs are still the best technique for detecting wildfire with our UNet providing an IoU of 93.58%, better than the baseline UNet by some 4.58%.

野火检测视觉变压器遥感影像深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。