arXiv:2605.00310cs.CVcs.AI2026-05被引 1

用下游任务评估遥感超分模型,发现传统指标失效。

Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration

论文配图:Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration
图 1 · 摘自论文原文
  • 构建多任务遥感超分评测集GeoSR-Bench,关联分辨率与实际应用。
  • 270组实验表明:清晰度提升不等于任务表现好,相关性甚至为负。
  • 适合关注遥感应用落地的研究者和开发者参考。

超分辨率(SR)技术在从低分辨率输入重建高分辨率图像方面取得显著进展,尤其在卫星遥感地球观测中广泛应用,涵盖城市规划、农业、生态和灾害响应等任务。然而,现有研究多依赖PSNR、SSIM等保真度指标,而真实价值在于支持地表监测等下游任务。为此,本文提出首个集成下游任务的超分评测基准GeoSR-Bench,包含约36,000个地理位置匹配、时间对齐且质量可控的图像对,覆盖500米至0.6米分辨率,涵盖多种地表类型。该基准直接将超分结果与土地覆盖分割、基础设施制图、生物物理变量估计等任务性能挂钩。我们在270个设置下评估了基于GAN、Transformer、神经算子和扩散模型的9种SR模型,涉及2类跨平台超分任务、3类下游任务模型和5类下游任务。结果表明,传统指标提升常无法带来任务性能增益,相关性甚至为负,说明现有指标难以指导实际应用模型选择,亟需将下游任务纳入超分模型开发与评估体系。

原文摘要 · Abstract (English)

Super-resolution (SR) techniques have made major advances in reconstructing high-resolution images from low-resolution inputs. The increased resolution provides visual enhancement and utility for monitoring tasks. In particular, SR has been increasingly developed for satellite-based Earth observation, with applications in urban planning, agriculture, ecology, and disaster response. However, existing SR studies and benchmarks typically use fidelity metrics such as PSNR or SSIM, whereas the true utility of super-resolved images lies in supporting downstream tasks such as land cover classification, biomass estimation, and change detection. To bridge this gap, we introduce GeoSR-Bench, a downstream task-integrated SR benchmark dataset to evaluate SR models beyond fidelity metrics. GeoSR-Bench comprises spatially co-located, temporally aligned, and quality-controlled image pairs from about 36,000 locations across diverse land covers, spanning resolutions from 500m to 0.6m. To the best of our knowledge, GeoSR-Bench is the first SR benchmark that directly connects improved image resolution from SR models with downstream Earth monitoring tasks, including land cover segmentation, infrastructure mapping, and biophysical variable estimation. Using GeoSR-Bench, we benchmark GAN, transformer, neural operator, and diffusion-based SR models on perceptual quality and downstream task performance. We conduct experiments with 270 settings, covering 2 cross-platform SR tasks, 9 SR models, 3 downstream task models, and 5 downstream tasks for each SR task. The results show that improvements in traditional SR metrics often do not correlate with gains in task performance, and the correlations can be negative, indicating that these metrics provide limited guidance for selecting superior models for downstream tasks. This reveals the need to integrate downstream tasks into SR model development and evaluation.

超分辨率遥感影像下游任务模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。