arXiv:2509.13289cs.CVeess.IV2025-09

用多模态特征评估生成图像真实感并定位不一致区域

Image Realness Assessment and Localization with Multimodal Features

  • 结合视觉语言模型生成文本描述,替代人工标注
  • 提升真实感预测准确率,生成密集真实感图
  • 适合需要图像质量评估与优化的AI生成场景

可靠量化AI生成图像的感知真实感并识别视觉不一致区域,对实际应用和通过真实感反馈改进生成模型至关重要。本文提出一种框架,利用在大规模数据集上训练的视觉语言模型生成的视觉不一致文本描述,实现对AI生成图像的整体客观真实感评估和局部不一致区域识别。实验表明,该多模态方法在客观真实感预测性能上有所提升,并生成了能有效区分真实与非真实空间区域的密集真实感图。

原文摘要 · Abstract (English)

A reliable method of quantifying the perceptual realness of AI-generated images and identifying visually inconsistent regions is crucial for practical use of AI-generated images and for improving photorealism of generative AI via realness feedback during training. This paper introduces a framework that accomplishes both overall objective realness assessment and local inconsistency identification of AI-generated images using textual descriptions of visual inconsistencies generated by vision-language models trained on large datasets that serve as reliable substitutes for human annotations. Our results demonstrate that the proposed multimodal approach improves objective realness prediction performance and produces dense realness maps that effectively distinguish between realistic and unrealistic spatial regions.

图像评估多模态真实感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。