检验合成遥感数据质量,发现主流指标失效,需结合人评与任务表现。
Benchmarking the Alignment of Data-Quality Metrics, Human Judgment and Land-Cover Segmentation Performance for Earth Observation
- 用生成模型合成遥感图像,对比自动指标、人类判断和下游任务表现。
- 旋转等细微变化使指标失真,但人眼识别不受影响,合成图反而更真实。
- 在土地覆盖分割任务中,低分合成数据能提升模型性能,适合实际应用。
大规模高质量数据对深度学习训练至关重要,但受限于获取成本与数据量。合成数据增强可通过生成模型生成逼真图像扩充数据集,其质量通常依赖FID、KID、IS、LPIPS、SSIM等视觉保真度指标评估。然而,这些指标(如广泛使用的FID)仅关注图像结构或分布相似性,不反映下游任务效用,在人类难以察觉的微小扰动下可能严重失准。本文系统评估了真实与生成的遥感数据集,比较自动指标、人类感知与下游语义分割任务表现。结果表明:语义保持的扰动(如旋转)显著改变指标分数,但人类识别几乎不变;部分自动指标得分较低的合成样本,其感知真实度却更高,且在混合真实-合成数据上训练的分割模型表现更优。研究表明,基于ImageNet预训练特征空间的质量指标无法可靠评估地理空间数据。因此,合成数据的质量评估应以任务表现和人类评价为基础。
原文摘要 · Abstract (English)
Volume and quality of datasets are crucial for deep learning model training, yet they are often constrained by availability and data acquisition costs. Synthetic data augmentation can extend existing datasets with realistic images, and the quality of these images is generally assessed through fidelity metrics such as FID, KID, IS, LPIPS and SSIM that measure structural or distributional similarity. However, such metrics, including the widely used FID, focus on visual fidelity without reflecting downstream utility, and can diverge from human perception under perturbations that are imperceptible to human observers. In this work, we systematically evaluate Earth observation datasets alongside synthetic counterparts generated by deep generative models, comparing automatic metrics against human perception and downstream tasks. Our results reveal a stark misalignment: semantics-preserving perturbations such as rotation drastically alter metric scores while leaving human recognition unaffected, and synthetic samples that score poorly on automatic metrics achieve comparable or higher perceived realism, and can improve downstream performance when combined with real data. By benchmarking semantic segmentation models trained on mixed real-synthetic datasets, we demonstrate that quality metrics rooted in ImageNet-pretrained feature spaces are unreliable indicators for geospatial data. Our findings underscore that automatic quality evaluation of synthetic datasets should be grounded in downstream task performance and human evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。