arXiv:2511.22812cs.CV2025-11

用文本生成图像并结合变形注意力,提升高分辨率土地覆盖分类精度。

LC4-DViT: Land-cover Creation for Land-cover Classification with Deformable Vision Transformer

论文配图:LC4-DViT: Land-cover Creation for Land-cover Classification with Deformable Vision Transformer
图 1 · 摘自论文原文
  • 用GPT-4o生成场景描述,合成类平衡的高清训练图像。
  • 在AID数据集上达到0.9572准确率,优于ViT等基线模型。
  • 适合需要高精度遥感分类的研究者与环境监测应用。

土地覆盖支撑生态系统服务、水文调控、灾害风险减缓和科学用地规划;及时准确的土地覆盖图对环境管理至关重要。基于遥感的土地覆盖分类虽具可扩展性,但仍受标注数据稀缺不平衡及高分辨率影像几何畸变的制约。本文提出LC4-DViT(基于可变形视觉变压器的土地覆盖生成与分类框架),融合生成式数据增强与变形感知视觉变压器。通过GPT-4o生成的场景描述与超分辨样例,构建类平衡、高保真的训练图像;DViT将DCNv4可变形卷积主干与视觉变压器编码器结合,联合捕捉细粒度几何结构与全局上下文。在包含海滩、桥梁、沙漠、森林、山地、池塘、港口、河流八类的AID数据集上,取得0.9572总体准确率、0.9576宏平均F1分数与0.9510 Cohen's Kappa,显著优于纯ViT基线(0.9274 OA, 0.9300 macro F1, 0.9169 Kappa)及ResNet50、MobileNetV2、FlashInternImage。在三类SIRI-WHU子集(港口、池塘、河流)上实现0.9333总体准确率、0.9316宏平均F1与0.8989 Kappa,表明良好迁移能力。基于GPT-4o的评分系统评估Grad-CAM热图显示,DViT注意力与水文意义结构对齐最佳。结果表明,描述驱动的生成增强与变形感知变压器结合,是高分辨率土地覆盖制图的有力路径。

原文摘要 · Abstract (English)

Land-cover underpins ecosystem services, hydrologic regulation, disaster-risk reduction, and evidence-based land planning; timely, accurate land-cover maps are therefore critical for environmental stewardship. Remote sensing-based land-cover classification offers a scalable route to such maps but is hindered by scarce and imbalanced annotations and by geometric distortions in high-resolution scenes. We propose LC4-DViT (Land-cover Creation for Land-cover Classification with Deformable Vision Transformer), a framework that combines generative data creation with a deformation-aware Vision Transformer. A text-guided diffusion pipeline uses GPT-4o-generated scene descriptions and super-resolved exemplars to synthesize class-balanced, high-fidelity training images, while DViT couples a DCNv4 deformable convolutional backbone with a Vision Transformer encoder to jointly capture fine-scale geometry and global context. On eight classes from the Aerial Image Dataset (AID)-Beach, Bridge, Desert, Forest, Mountain, Pond, Port, and River-DViT achieves 0.9572 overall accuracy, 0.9576 macro F1-score, and 0.9510 Cohen' s Kappa, improving over a vanilla ViT baseline (0.9274 OA, 0.9300 macro F1, 0.9169 Kappa) and outperforming ResNet50, MobileNetV2, and FlashInternImage. Cross-dataset experiments on a three-class SIRI-WHU subset (Harbor, Pond, River) yield 0.9333 overall accuracy, 0.9316 macro F1, and 0.8989 Kappa, indicating good transferability. An LLM-based judge using GPT-4o to score Grad-CAM heatmaps further shows that DViT' s attention aligns best with hydrologically meaningful structures. These results suggest that description-driven generative augmentation combined with deformation-aware transformers is a promising approach for high-resolution land-cover mapping.

土地覆盖遥感生成模型视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。