首个遥感图像推理编辑基准,评估模型在时空因果推理上的能力。
RS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image Editing

- 构建三类遥感编辑任务:时序、因果与空间推理。
- 最强模型在严格标准下仅24.28%准确率,多数模型接近零分。
- 适合遥感智能生成、地理推理与多模态模型研究者参考。
遥感图像编辑旨在根据自然语言指令修改图像,同时保持地理规则和传感器观测特性。现有基准主要针对自然图像或通用视觉场景,难以全面捕捉遥感编辑所需的推理、区域控制与传感器一致性能力。为此,我们提出首个推理引导的遥感图像编辑基准 RS-RIE-Bench。该基准将任务分为三类:时序推理、因果推理与空间推理,分别刻画遥感场景中的时间演变、因果后果与空间成像一致性。评估协议涵盖三个维度:目标区域合理性、非目标区域保真度及图像质量一致性。通过交叉评判一致性和分层专家评审验证了基于多模态大模型的评估可行性。对八种开源与闭源图像编辑模型的系统评估显示,当前模型在推理引导的遥感编辑上仍存在明显局限。即使最强模型在严格联合满足准则下整体准确率也仅为24.28%,所有模型平均宽松联合4成功率为32.23%。因果推理与空间推理尤为困难,部分开源模型在某些类别接近零分。结果表明,RS-RIE-Bench能有效揭示当前模型在地理推理、区域控制与传感器一致性生成方面的不足,为未来遥感智能编辑模型提供标准化基准与明确研究方向。
原文摘要 · Abstract (English)
Remote sensing image editing aims to modify remote sensing images according to natural language instructions while preserving geographic rules and sensor observation characteristics. Existing benchmarks mainly target natural images or general visual scenes, and thus may not fully capture the reasoning, regional control, and sensor-consistency abilities required in remote sensing editing. To fill this gap, we introduce RS-RIE-Bench, the first benchmark for reasoning-guided remote sensing image editing. RS-RIE-Bench organizes tasks into three categories: temporal reasoning, causal reasoning, and spatial reasoning. These categories capture temporal evolution, causal consequence, and spatial imaging consistency in remote sensing scenes. The evaluation protocol covers three dimensions: target region plausibility, non-target region preservation, and image quality consistency. We further demonstrate the feasibility of MLLM-based evaluation through cross-judge consistency analysis and stratified expert review. Systematic evaluation on eight open-source and closed-source image editing models shows that current models still have clear limitations in reasoning-guided remote sensing editing. Even the strongest model achieves only 24.28\% overall accuracy under the strict joint-satisfaction criterion, while the mean relaxed joint-4 success rate across all eight models is 32.23\%. Causal reasoning and spatial reasoning remain especially challenging, and several open-source models are close to zero in some categories. These results show that RS-RIE-Bench can effectively reveal the limitations of current models in geographic reasoning, regional control, and sensor-consistent generation. It also provides a standardized benchmark and a clear research direction for future remote sensing intelligent editing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。