通过动态分辨率与多尺度对齐,提升遥感图像多模态理解精度与效率。
Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism

- 根据图像内容复杂度动态调整分辨率,节省计算资源。
- 在物体、局部和全局三层实现跨模态语义对齐,减少错配。
- 适合遥感智能解译、环境监测等需要高精度图文匹配的场景。
遥感图像多模态融合是克服单一数据源局限、提升地表信息提取精度的核心技术,在环境监测和城市规划等领域具有重要应用价值。针对现有方法存在固定分辨率难以兼顾效率与细节、单尺度对齐缺乏语义层次的问题,本文提出集成两项创新的视觉-语言模型(VLM)框架:动态分辨率输入策略(DRIS)与多尺度视觉-语言对齐机制(MS-VLAM)。DRIS采用从粗到细的自适应策略,根据图像内容复杂度分配计算资源,有效保留关键细粒度特征并降低冗余开销。MS-VLAM构建覆盖物体、局部区域与全局层面的三级对齐机制,系统捕捉跨模态语义一致性,缓解语义错位与粒度不平衡问题。在RS-GPT4V数据集上的实验表明,该框架在图像描述生成和跨模态检索任务中显著提升语义理解准确率与计算效率。相比传统方法,在图像描述生成的BLEU-4与CIDEr指标,以及跨模态检索的R@10指标上均取得更优表现。该技术框架为构建高效、鲁棒的多模态遥感系统提供了新路径,奠定了理论基础并为智能遥感解译的工程应用提供指导。
原文摘要 · Abstract (English)
Multimodal fusion of remote sensing images serves as a core technology for overcoming the limitations of single-source data and improving the accuracy of surface information extraction, which exhibits significant application value in fields such as environmental monitoring and urban planning. To address the deficiencies of existing methods, including the failure of fixed resolutions to balance efficiency and detail, as well as the lack of semantic hierarchy in single-scale alignment, this study proposes a Vision-language Model (VLM) framework integrated with two key innovations: the Dynamic Resolution Input Strategy (DRIS) and the Multi-scale Vision-language Alignment Mechanism (MS-VLAM).Specifically, the DRIS adopts a coarse-to-fine approach to adaptively allocate computational resources according to the complexity of image content, thereby preserving key fine-grained features while reducing redundant computational overhead. The MS-VLAM constructs a three-tier alignment mechanism covering object, local-region and global levels, which systematically captures cross-modal semantic consistency and alleviates issues of semantic misalignment and granularity imbalance.Experimental results on the RS-GPT4V dataset demonstrate that the proposed framework significantly improves the accuracy of semantic understanding and computational efficiency in tasks including image captioning and cross-modal retrieval. Compared with conventional methods, it achieves superior performance in evaluation metrics such as BLEU-4 and CIDEr for image captioning, as well as R@10 for cross-modal retrieval. This technical framework provides a novel approach for constructing efficient and robust multimodal remote sensing systems, laying a theoretical foundation and offering technical guidance for the engineering application of intelligent remote sensing interpretation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。