用可扩展框架评估生成图像的逼真度,发现生成模型在模拟恶劣天气上优于传统方法。
Scalable Evaluation of the Realism of Synthetic Environmental Augmentations in Images
- 构建视觉-语言模型与嵌入分布分析结合的评估框架
- 生成模型在雾天等场景下接受率是规则方法的3.6倍
- 适用于自动驾驶测试数据生成,尤其适合缺乏真实数据的场景
AI系统评估常需合成测试样本,尤其对罕见或安全关键情形。生成式AI可通过可控图像编辑生成此类数据,但其有效性取决于图像逼真度。本文提出一个可扩展框架,评估将雾、雨、雪和夜间条件添加至车载摄像头图像的合成方法。基于40张晴天图像,对比规则增强库与生成式图像编辑模型。使用两种互补自动化指标:视觉-语言模型(VLM)裁判评估感知逼真度,嵌入分布分析衡量与真实恶劣条件图像的相似性。生成式方法显著优于规则方法,最佳生成模型的接受率约为最优规则方法的3.6倍。性能因条件而异:雾最易模拟,夜间仍具挑战。值得注意的是,即使真实恶劣条件图像也未获完全接受,确立了实际评判上限。在此标准下,领先生成方法在多数条件下达到或超过真实图像表现。结果表明,现代生成式图像编辑模型可实现大规模真实感恶劣条件图像生成,支持评估流程。本框架为可扩展逼真度评估提供实用方案,未来仍需人类实验验证。
原文摘要 · Abstract (English)
Evaluation of AI systems often requires synthetic test cases, particularly for rare or safety-critical conditions that are difficult to observe in operational data. Generative AI offers a promising approach for producing such data through controllable image editing, but its usefulness depends on whether the resulting images are sufficiently realistic to support meaningful evaluation. We present a scalable framework for assessing the realism of synthetic image-editing methods and apply it to the task of adding environmental conditions-fog, rain, snow, and nighttime-to car-mounted camera images. Using 40 clear-day images, we compare rule-based augmentation libraries with generative AI image-editing models. Realism is evaluated using two complementary automated metrics: a vision-language model (VLM) jury for perceptual realism assessment, and embedding-based distributional analysis to measure similarity to genuine adverse-condition imagery. Generative AI methods substantially outperform rule-based approaches, with the best generative method achieving approximately 3.6 times the acceptance rate of the best rule-based method. Performance varies across conditions: fog proves easiest to simulate, while nighttime transformations remain challenging. Notably, the VLM jury assigns imperfect acceptance even to real adverse-condition imagery, establishing practical ceilings against which synthetic methods can be judged. By this standard, leading generative methods match or exceed real-image performance for most conditions. These results suggest that modern generative image-editing models can enable scalable generation of realistic adverse-condition imagery for evaluation pipelines. Our framework therefore provides a practical approach for scalable realism evaluation, though validation against human studies remains an important direction for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。