构建首个大规模航拍图像指代表达分割数据集,支持从现代到历史影像的精准语义定位。
Generalized Referring Expression Segmentation on Aerial Photos

- 通过规则生成+大模型增强自动构建航拍指代表达数据集
- 包含3.7万张图、150万条表达,覆盖21类目标且可处理像素级细小目标
- 模型在历史影像降质条件下仍保持高精度,适合遥感与考古应用
指代表达分割是计算机视觉中融合自然语言理解与视觉精确定位的基础任务。针对航拍图像(如无人机拍摄、历史航空档案、高分辨率卫星影像等)存在分辨率差异大、色彩不一致、目标缩小至数像素、场景物体密度高且常部分遮挡等问题,本文提出Aerial-D,一个大规模航拍图像指代表达分割数据集。该数据集包含37,288张图像、1,522,523条指代表达,标注259,709个目标,涵盖21种类别,包括车辆、基础设施及地表覆盖类型。数据集通过全自动化流程构建,结合规则生成与大语言模型(LLM)增强,提升表达多样性与视觉细节聚焦;并加入滤镜模拟历史影像条件。采用RSRefSeg架构,在Aerial-D与已有航拍数据集联合训练,实现对现代与历史影像的统一实例与语义分割。实验表明,联合训练在当前基准上表现优异,且在单色、棕褐色、颗粒化等历史影像退化条件下仍保持高准确率。数据集、训练模型及完整软件管道已公开:https://luispl77.github.io/aerial-d。
原文摘要 · Abstract (English)
Referring expression segmentation is a fundamental task in computer vision that integrates natural language understanding with precise visual localization of target regions. Considering aerial imagery (e.g., modern aerial photos collected through drones, historical photos from aerial archives, high-resolution satellite imagery, etc.) presents unique challenges because spatial resolution varies widely across datasets, the use of color is not consistent, targets often shrink to only a few pixels, and scenes contain very high object densities and objects with partial occlusions. This work presents Aerial-D, a new large-scale referring expression segmentation dataset for aerial imagery, comprising 37,288 images with 1,522,523 referring expressions that cover 259,709 annotated targets, spanning across individual object instances, groups of instances, and semantic regions covering 21 distinct classes that range from vehicles and infrastructure to land coverage types. The dataset was constructed through a fully automatic pipeline that combines systematic rule-based expression generation with a Large Language Model (LLM) enhancement procedure that enriched both the linguistic variety and the focus on visual details within the referring expressions. Filters were additionally used to simulate historic imaging conditions for each scene. We adopted the RSRefSeg architecture, and trained models on Aerial-D together with prior aerial datasets, yielding unified instance and semantic segmentation from text for both modern and historical images. Results show that the combined training achieves competitive performance on contemporary benchmarks, while maintaining strong accuracy under monochrome, sepia, and grainy degradations that appear in archival aerial photography. The dataset, trained models, and complete software pipeline are publicly available at https://luispl77.github.io/aerial-d .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。