arXiv:2409.01138cs.CVcs.AI2024-09被引 5

用真实数据训练生成罕见卫星图像,验证了生成效果与人类判断的不一致。

Generating Synthetic Satellite Imagery for Rare Objects: An Empirical Comparison of Models and Metrics

  • 微调生成模型,通过文本或游戏引擎布局图生成卫星图像
  • 即使仅有约400个真实样本,也能生成逼真图像
  • 自动指标与人类评分呈强负相关,提示需谨慎依赖量化评估

生成式深度学习架构可生成高分辨率逼真假图像,带来重大社会影响。关键问题是:在小众领域中,生成逼真图像有多难?实现特定图像内容所需的迭代过程难以自动化和控制。尤其对稀有类别,仍难以评估生成图像的真实性(是否逼真)与可控性(能否被人类输入有效引导)。本文对微调后的生成架构进行大规模实证评估,聚焦核电厂这一罕见对象类别——全球仅约400座,典型代表了训练与测试数据受限的诸多场景。通过文本输入或来自游戏引擎的图像输入(可精确指定建筑布局),生成合成卫星图像。使用常见自动评估指标,并结合用户研究的人类判断,评估图像可信度。结果表明,即使对于稀有对象,基于文本或详细布局生成真实感卫星图像仍是可行的。与以往研究一致,发现自动指标常与人类感知不一致,甚至呈现显著负相关。

原文摘要 · Abstract (English)

Generative deep learning architectures can produce realistic, high-resolution fake imagery -- with potentially drastic societal implications. A key question in this context is: How easy is it to generate realistic imagery, in particular for niche domains. The iterative process required to achieve specific image content is difficult to automate and control. Especially for rare classes, it remains difficult to assess fidelity, meaning whether generative approaches produce realistic imagery and alignment, meaning how (well) the generation can be guided by human input. In this work, we present a large-scale empirical evaluation of generative architectures which we fine-tuned to generate synthetic satellite imagery. We focus on nuclear power plants as an example of a rare object category - as there are only around 400 facilities worldwide, this restriction is exemplary for many other scenarios in which training and test data is limited by the restricted number of occurrences of real-world examples. We generate synthetic imagery by conditioning on two kinds of modalities, textual input and image input obtained from a game engine that allows for detailed specification of the building layout. The generated images are assessed by commonly used metrics for automatic evaluation and then compared with human judgement from our conducted user studies to assess their trustworthiness. Our results demonstrate that even for rare objects, generation of authentic synthetic satellite imagery with textual or detailed building layouts is feasible. In line with previous work, we find that automated metrics are often not aligned with human perception -- in fact, we find strong negative correlations between commonly used image quality metrics and human ratings.

生成模型卫星图像稀有对象人类评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。