arXiv:2608.13344cs.AI2026-08

构建首个长时序遥感多阶段推理基准,提升模型对地理演变的连贯理解能力。

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

论文配图:LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning
图 1 · 摘自论文原文
  • 设计包含120万样本的长序列遥感问答数据集,平均15.14帧,最长30帧。
  • 提出结构化思维链监督与多维度奖励机制,显著提升模型长程推理性能。
  • 在12项长时序任务上全面领先,适用于遥感智能分析、环境监测等场景。

长时序地球观测推理要求模型能够组织多阶段地理演化过程,定位空间变化,检测时间异常,并从扩展图像序列中推断未来。然而,现有遥感视觉语言模型主要关注单张图像、图像对或短序列,难以在相关帧和区域中实现可靠定位。我们引入LongEarth-Bench,一个包含约120,000个问答样本的数据集,源自117,000张唯一图像,其序列平均15.14帧,最长可达30帧,涵盖12项任务:演化总结、空间推理、异常识别与逻辑预测。其中30,000个样本提供结构化推理轨迹,明确关联关键帧与变化区域至最终答案。我们通过显式序列标识符与结构化思维链监督进行有监督微调,构建LongEarth。在此基础上,LongEarth-R1采用组相对策略优化,结合格式、时间与空间奖励。LongEarth-R1在所有12项长序列任务中表现最优,同时在标准遥感基准上保持竞争力。

原文摘要 · Abstract (English)

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.

遥感推理长序列理解视觉语言模型多阶段推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。