arXiv:2603.14941cs.AI2026-03被引 3

用统一模型同时理解遥感变化和预测未来,效果远超更大模型。

RS-WorldModel: a Unified Model for Remote Sensing Understanding and Future Sense Forecasting

  • 联合训练理解与预测任务,共享时空先验知识
  • 20亿参数模型在问答任务上超越120倍大的开源模型
  • 支持文本引导的未来场景生成,性能优于闭源模型

遥感世界模型旨在解释观测到的变化并预测可能的未来,这两项任务共享时空先验。现有方法通常分别处理,限制了跨任务迁移。我们提出RS-WorldModel,一个统一的遥感世界模型,可联合处理时空变化理解与文本引导的未来场景预测,并构建了包含110万样本、富含语言标注的RSWBench-1.1M数据集。该模型分三阶段训练:(1) 地理感知生成预训练(GAGP),以地理和获取元数据为条件进行预测;(2) 协同指令微调(SIT),联合训练理解与预测;(3) 可验证强化优化(VRO),使用可验证的任务特定奖励优化输出。仅20亿参数的模型在多数时空变化问答指标上超越最大达120倍的开源模型。在文本引导未来场景生成上,FID达43.13,优于所有开源基线及闭源模型Gemini-2.5-Flash Image (Nano Banana)。

原文摘要 · Abstract (English)

Remote sensing world models aim to both explain observed changes and forecast plausible futures, two tasks that share spatiotemporal priors. Existing methods, however, typically address them separately, limiting cross-task transfer. We present RS-WorldModel, a unified world model for remote sensing that jointly handles spatiotemporal change understanding and text-guided future scene forecasting, and we build RSWBench-1.1M, a 1.1 million sample dataset with rich language annotations covering both tasks. RS-WorldModel is trained in three stages: (1) Geo-Aware Generative Pre-training (GAGP) conditions forecasting on geographic and acquisition metadata; (2) synergistic instruction tuning (SIT) jointly trains understanding and forecasting; (3) verifiable reinforcement optimization (VRO) refines outputs with verifiable, task-specific rewards. With only 2B parameters, RS-WorldModel surpasses open-source models up to 120$ \times $ larger on most spatiotemporal change question-answering metrics. It achieves an FID of 43.13 on text-guided future scene forecasting, outperforming all open-source baselines as well as the closed-source Gemini-2.5-Flash Image (Nano Banana).

遥感世界模型未来预测多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。