用卫星图精准识别人群,突破传统图像局限
Crowd Detection Using Very-Fine-Resolution Satellite Imagery
- 基于点的神经网络,融合场景与个体特征提升识别能力
- 在12万+标注人群数据上达66.12%最高F1分数
- 首个专用于卫星人群检测的数据集,适合遥感与安防研究
精准人群检测对公共安全与历史模式分析至关重要,但现有方法依赖地面和航拍图像,时空覆盖有限。随着超精细分辨率(约0.3米)卫星影像的发展,大规模人群活动分析迎来新机遇,但该领域尚未被充分探索。为此,本文提出CrowdSat-Net——一种基于点的卷积神经网络,包含两个创新模块:双上下文渐进注意力网络(DCPAN)通过聚合场景上下文与局部个体特征增强表征;高频引导可变形上采样器(HFGDU)利用频域引导的可变形卷积恢复上采样过程中的高频信息。为验证模型有效性,我们构建了首个专为人群检测设计的超精细卫星影像数据集CrowdSat,涵盖来自北京三号N、吉林一号、高分四号A及Google Earth的多源平台数据,共超过12万个人工标注样本。实验表明,CrowdSat-Net在该数据集上达到66.12%的最高F1分数与73.23%的精度,优于五种主流点基检测方法,分别领先1.71%与2.42%。消融实验证明了DCPAN与HFGDU模块的关键作用,跨区域评估进一步验证了模型的空间泛化能力。本研究同时提供了新的网络架构与基准数据集,推动人群检测技术发展。
原文摘要 · Abstract (English)
Accurate crowd detection (CD) is critical for public safety and historical pattern analysis, yet existing methods relying on ground and aerial imagery suffer from limited spatio-temporal coverage. The development of very-fine-resolution (VFR) satellite sensor imagery (e.g., ~0.3 m spatial resolution) provides unprecedented opportunities for large-scale crowd activity analysis, but it has never been considered for this task. To address this gap, we proposed CrowdSat-Net, a novel point-based convolutional neural network, which features two innovative components: Dual-Context Progressive Attention Network (DCPAN) to improve feature representation of individuals by aggregating scene context and local individual characteristics, and High-Frequency Guided Deformable Upsampler (HFGDU) that recovers high-frequency information during upsampling through frequency-domain guided deformable convolutions. To validate the effectiveness of CrowdSat-Net, we developed CrowdSat, the first VFR satellite imagery dataset designed specifically for CD tasks, comprising over 120k manually labeled individuals from multi-source satellite platforms (Beijing-3N, Jilin-1 Gaofen-04A and Google Earth) across China. In the experiments, CrowdSat-Net was compared with five state-of-the-art point-based CD methods (originally designed for ground or aerial imagery) using CrowdSat and achieved the largest F1-score of 66.12% and Precision of 73.23%, surpassing the second-best method by 1.71% and 2.42%, respectively. Moreover, extensive ablation experiments validated the importance of the DCPAN and HFGDU modules. Furthermore, cross-regional evaluation further demonstrated the spatial generalizability of CrowdSat-Net. This research advances CD capability by providing both a newly developed network architecture for CD and a pioneering benchmark dataset to facilitate future CD development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。