用框选区域引导模型,精准理解遥感图像中的多模态信息。
Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- 用户只需框选感兴趣区域,模型自动生成匹配的分割图和描述。
- 引入关系感知解码器与跨模态对齐损失,提升相似目标识别精度。
- 适合遥感图像分析、智能标注等需要精准语义理解的场景。
近期图像理解进展使得大语言模型能用于遥感领域的多模态推理,但现有方法在仅提供简单通用文本提示时,难以引导模型聚焦用户关注区域。此外,在大规模航空影像中,许多物体视觉外观高度相似且存在丰富互相关系,进一步增加了准确识别的难度。为此,本文提出面向视觉提示引导的多模态遥感图像理解的跨模态上下文感知学习框架(CLV-Net)。CLV-Net允许用户通过一个简单的边界框指定兴趣区域,并据此引导模型生成符合用户意图的分割掩码与文字描述。核心设计包括:上下文感知掩码解码器,用于建模并整合对象间关系以增强目标表征;以及语义与关系对齐模块:跨模态语义一致性损失提升对视觉相似目标的细粒度区分能力,关系一致性损失确保文本描述中的关系与图像中的视觉交互保持一致。在两个基准数据集上的全面实验表明,CLV-Net优于现有方法,建立了新的最先进性能。模型能有效捕捉用户意图,生成精确且与意图一致的多模态输出。
原文摘要 · Abstract (English)
Recent advances in image understanding have enabled methods that leverage large language models for multimodal reasoning in remote sensing. However, existing approaches still struggle to steer models to the user-relevant regions when only simple, generic text prompts are available. Moreover, in large-scale aerial imagery many objects exhibit highly similar visual appearances and carry rich inter-object relationships, which further complicates accurate recognition. To address these challenges, we propose Cross-modal Context-aware Learning for Visual Prompt-Guided Multimodal Image Understanding (CLV-Net). CLV-Net lets users supply a simple visual cue, a bounding box, to indicate a region of interest, and uses that cue to guide the model to generate correlated segmentation masks and captions that faithfully reflect user intent. Central to our design is a Context-Aware Mask Decoder that models and integrates inter-object relationships to strengthen target representations and improve mask quality. In addition, we introduce a Semantic and Relationship Alignment module: a Cross-modal Semantic Consistency Loss enhances fine-grained discrimination among visually similar targets, while a Relationship Consistency Loss enforces alignment between textual relations and visual interactions. Comprehensive experiments on two benchmark datasets show that CLV-Net outperforms existing methods and establishes new state-of-the-art results. The model effectively captures user intent and produces precise, intention-aligned multimodal outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。