arXiv:2510.21821cs.CVcs.AI2025-10

测试GPT-4o/DALL-E3生成图像对提示词的忠实度,发现有近16%属性出错。

Prompt fidelity of ChatGPT4o / Dall-E3 text-to-image visualisations

  • 用自动生成提示词,测试图像是否准确呈现指定属性
  • 整体属性错误率15.6%,年龄描述出错最严重
  • 适合关注生成模型偏差与评估的研究者

本研究通过分析自动生成提示词中明确指定的属性在图像中是否被正确呈现,评估了ChatGPT4o/DALL-E3文本到图像生成的提示词忠实度。使用两个公开数据集,分别包含200张文化与创意产业女性工作者和230张博物馆策展人图像,评估了个人属性(年龄、发型)、外貌特征(着装、眼镜)及随身物品(姓名牌、笔记本)的准确性。尽管多数属性正确呈现,但DALL-E3在所有属性中的错误率为15.6%(共710个属性),其中随身物品错误最少,个人外貌中等,人物自身描绘错误最高,尤其是年龄。结果表明生成模型存在可量化的提示到图像忠实度差距,对偏见检测与模型评估具有重要启示。

原文摘要 · Abstract (English)

This study examines the prompt fidelity of ChatGPT4o / DALL-E3 text-to-image visualisations by analysing whether attributes explicitly specified in autogenously generated prompts are correctly rendered in the resulting images. Using two public-domain datasets comprising 200 visualisations of women working in the cultural and creative industries and 230 visualisations of museum curators, the study assessed accuracy across personal attributes (age, hair), appearance (attire, glasses), and paraphernalia (name tags, clipboards). While correctly rendered in most cases, DALL-E3 deviated from prompt specifications in 15.6% of all attributes (n=710). Errors were lowest for paraphernalia, moderate for personal appearance, and highest for depictions of the person themselves, particularly age. These findings demonstrate measurable prompt-to-image fidelity gaps with implications for bias detection and model evaluation.

文本生成图像提示词忠实度模型评估偏差检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。