用视觉变压器提升遥感图像分类泛化能力
How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
- 结合视觉变换器与多任务学习,增强空间感知与上下文理解
- 在多个遥感数据集上实现更高准确率和更强跨域适应性
- 适合遥感图像分析、地理信息研究等领域的研究人员
遥感数据集为土地利用分类、目标存在检测和城乡分类等关键任务提供了巨大潜力。然而,现有研究多聚焦于特定任务或数据集,限制了其在各类遥感分类挑战中的泛化能力。为此,我们提出一种新模型 SpatialNet-ViT,融合视觉变换器(ViTs)与多任务学习(MTL),结合空间感知与上下文理解,显著提升分类准确率与可扩展性。同时,通过数据增强、迁移学习和多任务学习等技术,进一步增强模型鲁棒性与跨多样化数据集的泛化能力。
原文摘要 · Abstract (English)
Remote sensing datasets offer significant promise for tackling key classification tasks such as land-use categorization, object presence detection, and rural/urban classification. However, many existing studies tend to focus on narrow tasks or datasets, which limits their ability to generalize across various remote sensing classification challenges. To overcome this, we propose a novel model, SpatialNet-ViT, leveraging the power of Vision Transformers (ViTs) and Multi-Task Learning (MTL). This integrated approach combines spatial awareness with contextual understanding, improving both classification accuracy and scalability. Additionally, techniques like data augmentation, transfer learning, and multi-task learning are employed to enhance model robustness and its ability to generalize across diverse datasets
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。