GeoMag让遥感图像解析更精准高效,支持像素级细节识别。
GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing
- 根据提示语动态调整关注区域,智能聚焦关键目标
- 在10个基准上实现像素级解析领先,计算成本降低显著
- 适合需要高精度遥感分析的科研与工程人员
视觉语言模型(VLM)在遥感(RS)图像理解中取得进展,具备识别和描述地理实体的能力。然而现有RS-VLM多局限于图像级和区域级任务,难以处理像素级任务,且在小目标识别中表现不佳。同时,处理高分辨率遥感图像时计算开销大,限制了实际应用。为此,我们提出GeoMag(地理放大器),一种端到端通用大模型框架。GeoMag基于提示语语义动态调整注意力范围,实现多粒度遥感图像解析。该方法引入任务驱动的多粒度分辨率调节(TMRA)和提示引导的语义感知裁剪(PSC),自适应降低无关区域的空间分辨率,增强相关区域的视觉表征。该策略提升了对关键目标区域的感知能力,抑制背景冗余,降低高分辨率遥感图像解析的计算成本。在10个基准上的广泛对比实验表明,GeoMag不仅在像素级任务上表现优异,且在其他粒度任务上也保持竞争力。
原文摘要 · Abstract (English)
The application of Vision-Language Models (VLMs) in remote sensing (RS) image understanding has achieved notable progress, demonstrating the basic ability to recognize and describe geographical entities. However, existing RS-VLMs are mostly limited to image-level and region-level tasks, lacking the capability to handle pixel-level tasks and performing poorly in small-object recognition scenarios. Moreover, RS-VLMs consume significant computational resources when processing high-resolution RS images, further restricting their practical applicability. In this context, we propose GeoMag (Geographical Magnifier), an end-to-end general-purpose large model framework for RS. GeoMag dynamically focuses the attention scope based on prompt semantics to effectively perform remote sensing image parsing across multiple levels of granularity. This method introduces Task-driven Multi-granularity Resolution Adjustment (TMRA) and Prompt-guided Semantic-aware Cropping (PSC), which adaptively reduce the spatial resolution of task-irrelevant regions while enhancing the visual representation of task-relevant areas. This approach improves the model's perception of critical target regions, suppresses background redundancy, and reduces the computational cost of interpreting high-resolution RS imagery. Extensive comparative experiments on 10 benchmarks demonstrate that GeoMag not only excels in handling pixel-level tasks but also maintains competitive performance across tasks of other granularities compared to existing RS-VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。