专为遥感设计的物体级多模态大模型,提升目标定位与属性描述精度。
EagleVision: Object-level Attribute Multimodal LLM for Remote Sensing
- 引入属性解耦模块,分离视觉特征以精准表达物体属性。
- 在95000+样本数据集上训练,实现遥感图像中细粒度目标检测与属性理解领先性能。
- 适合遥感智能分析、地理信息挖掘等需要精确物体理解的场景。
近年来,多模态大语言模型(MLLMs)在多种视觉任务中表现优异。然而,在遥感(RS)领域,高分辨率图像和小目标占比导致现有模型难以胜任以物体为中心的任务,尤其在精确定位与细粒度属性描述方面表现不佳。当前的遥感MLLMs仍逊于传统视觉感知模型,仅提供粗粒度图像理解,限制了其在真实场景中的应用。为此,我们提出EagleVision,一种专为遥感设计的物体级多模态大模型,在目标检测与属性理解方面表现出色。通过引入属性解耦模块,模型学习解耦的视觉标记来表达不同属性。为支持物体级视觉-语言对齐,我们构建了首个大规模遥感物体属性理解数据集EVAttrs-95K,并设计了新型评估基准EVBench。EagleVision在细粒度目标检测与物体属性理解任务上均达到当前最优水平,凸显了检测与理解能力之间的相互促进作用。代码、模型、数据及演示将开源。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLMs) have demonstrated impressive results in various visual tasks. However, in remote sensing (RS), high resolution and small proportion of objects pose challenges to existing MLLMs, which struggle with object-centric tasks, particularly in precise localization and fine-grained attribute description for each object. These RS MLLMs have not yet surpassed classical visual perception models, as they only provide coarse image understanding, leading to limited gains in real-world scenarios. To address this gap, we establish EagleVision, an MLLM tailored for remote sensing that excels in object detection and attribute comprehension. Equipped with the Attribute Disentangle module, EagleVision learns disentanglement vision tokens to express distinct attributes. To support object-level visual-language alignment, we construct EVAttrs-95K, the first large-scale object attribute understanding dataset in RS for instruction tuning, along with a novel evaluation benchmark, EVBench. EagleVision achieves state-of-the-art performance on both fine-grained object detection and object attribute understanding tasks, highlighting the mutual promotion between detection and understanding capabilities in MLLMs. The code, model, data, and demo will be available at https://github.com/XiangTodayEatsWhat/EagleVision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。