综述遥感视觉与多模态大模型研究进展与挑战
A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
- 系统梳理遥感视觉与多模态大模型的架构与训练方法
- 指出数据对齐、跨模态迁移与可扩展性是主要瓶颈
- 适合遥感智能分析、地理信息科学方向研究人员参考
遥感基础模型,尤其是视觉与多模态模型的快速发展,显著提升了智能地理空间数据解析能力。这些模型融合光学、雷达、LiDAR影像与文本、地理信息等多源数据,实现对遥感数据更全面的分析,提升目标检测、地表覆盖分类、变化检测等任务性能。然而,数据类型多样、大规模标注数据稀缺、多模态融合复杂等问题仍制约其应用。此外,训练与微调多模态模型的高算力需求也增加了实际部署难度。本文综述遥感视觉与多模态基础模型的最新进展,涵盖架构设计、训练方法、数据集与应用场景,讨论数据对齐、跨模态迁移学习及可扩展性等关键挑战,并提出未来研究方向。附录资源列表见:https://github.com/IRIP-BUAA/A-Review-for-remote-sensing-vision-language-models。
原文摘要 · Abstract (English)
The rapid advancement of remote sensing foundation models, particularly vision and multimodal models, has significantly enhanced the capabilities of intelligent geospatial data interpretation. These models combine various data modalities, such as optical, radar, and LiDAR imagery, with textual and geographic information, enabling more comprehensive analysis and understanding of remote sensing data. The integration of multiple modalities allows for improved performance in tasks like object detection, land cover classification, and change detection, which are often challenged by the complex and heterogeneous nature of remote sensing data. However, despite these advancements, several challenges remain. The diversity in data types, the need for large-scale annotated datasets, and the complexity of multimodal fusion techniques pose significant obstacles to the effective deployment of these models. Moreover, the computational demands of training and fine-tuning multimodal models require significant resources, further complicating their practical application in remote sensing image interpretation tasks. This paper provides a comprehensive review of the state-of-the-art in vision and multimodal foundation models for remote sensing, focusing on their architecture, training methods, datasets and application scenarios. We discuss the key challenges these models face, such as data alignment, cross-modal transfer learning, and scalability, while also identifying emerging research directions aimed at overcoming these limitations. Our goal is to provide a clear understanding of the current landscape of remote sensing foundation models and inspire future research that can push the boundaries of what these models can achieve in real-world applications. The list of resources collected by the paper can be found in the https://github.com/IRIP-BUAA/A-Review-for-remote-sensing-vision-language-models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。