arXiv:2606.10819cs.CVcs.AI2026-06被引 3

将遥感多模态大模型扩展至6种传感器和9类任务,统一处理跨模态地理信息。

Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks

论文配图:Earth-OneVision: Extending Remote Sensing Multimodal Large Language Models to More Sensor Modalities and Tasks
图 1 · 摘自论文原文
  • 设计三种机制解决多模态对齐、空间输出统一与域差异分步适应问题。
  • 在6种传感器上实现87.5%光学定位准确率,超越7B模型超7个百分点。
  • 适合遥感、地理信息、多源数据融合等领域的研究人员使用。

遥感多模态大模型(RS-MLLMs)可实现对地球观测图像的自然语言理解与空间推理,但现有模型仅支持有限传感器类型和任务,导致地球认知碎片化,跨模态地学知识未被充分挖掘。本文提出Earth-OneVision,一个20亿参数的统一框架,涵盖光学、合成孔径雷达(SAR)、红外、多光谱、时序与视频共六种传感器模态,并覆盖9类任务。通过三个专用机制:全粒度视觉-语言对齐(FGVLA)实现多层次视觉特征与多维语言空间的对齐;空间-语言同构序列化(SLIS)将异构空间输出转化为自回归标记;渐进式跨模态适配(PCMA)将复合域差距分解为视角与成像物理差异的逐步处理。为支持联合训练,构建了包含约3400万条问答对的MMRS-OneVision数据集,覆盖所有六种传感器及跨传感器融合任务,显著超过现有遥感多模态指令数据集。尽管仅有20亿参数,Earth-OneVision在多个基准测试中表现优异,性能媲美甚至超越40亿至720亿参数的模型。其在OPT-RSVG测试集上光学视觉定位准确率达87.52%([email protected]),在SAR VQA基准SARLANG-Bench上达80.68%,优于70亿模型超7%;在BigEarthNet-MS多光谱分类测试集上召回率达75.74%;在EarthMind-Bench跨模态推理测试中,多项选择题准确率为81.94%。

原文摘要 · Abstract (English)

RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery. However, existing models support only a narrow range of sensor types and tasks, yielding a fragmented view of the earth and leaving cross-modal geoscientific knowledge largely unexploited. This work presents Earth-OneVision, a 2B RS-MLLM that unifies six sensor modalities (i.e., optical, SAR, infrared, multispectral, temporal, and video) and cross-sensor fusion across 9 task categories within a single autoregressive framework. Three dedicated mechanisms address three bottlenecks. Full-Granularity Vision-Language Alignment (FGVLA) aligns multi-level visual features with the multi-dimensional language space. Spatial-Linguistic Isomorphic Serialization (SLIS) unifies heterogeneous spatial outputs as autoregressive tokens. Progressive Cross-Modality Adaptation (PCMA) decomposes the compound domain gap into sequential stages, tackling the viewpoint and imaging physics gaps in turn. To support joint training, MMRS-OneVision is constructed with ~34M QA pairs spanning all six sensor modalities and cross-sensor fusion across 9 task categories, substantially exceeding existing RS multimodal instruction datasets. With only 2B parameters, Earth-OneVision achieves competitive or state-of-the-art results across extensive benchmarks, consistently matching or outperforming 4B-72B RS-MLLMs. It achieves 87.52% [email protected] on the OPT-RSVG testset for optical visual grounding and 80.68% on the SAR VQA benchmark SARLANG-Bench, exceeding 7B models by over 7%. It further achieves 75.74% recall on the BigEarthNet-MS testset for multispectral classification, and 81.94% MCQ accuracy on EarthMind-Bench for cross-modality reasoning.

遥感多模态大模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。