评测扩散模型生成历史图像的准确性,发现存在刻板印象和时代错乱。
Synthetic History: Evaluating Visual Representations of the Past in Diffusion Models
- 构建3万张合成历史图像数据集,用精准提示词覆盖多时期人类活动。
- 发现模型常混淆时代风格、引入现代物品、性别种族分布失真。
- 提供可复现评估方案,适合关注历史准确性与生成伦理的研究者。
随着文本到图像扩散模型在内容创作中日益重要,其社会与文化影响受到越来越多关注。尽管已有研究聚焦于人口与文化偏见,但这些模型对历史语境的准确表现仍缺乏系统探讨。为此,我们提出一个评估历史语境表现的基准:结合由三款先进扩散模型基于精心设计提示词生成的30,000张合成图像构成的HistVis数据集,以及可复现的评估协议。从三个方面评估生成图像:(1) 隐含风格关联性:分析特定历史时期默认视觉风格;(2) 历史一致性:识别如现代物品出现在前现代社会背景中的时代错乱;(3) 人口代表性:将生成的种族与性别分布与历史合理基准进行对比。结果揭示,生成的历史图像普遍存在系统性偏差,模型常通过未明示的风格线索刻板化过去时代,引入时代错乱,并未能反映合理的种族与性别分布。本研究通过提供可复现的评估基准,为构建更具历史准确性的文本到图像模型迈出第一步。
原文摘要 · Abstract (English)
As Text-to-Image (TTI) diffusion models become increasingly influential in content creation, growing attention is being directed toward their societal and cultural implications. While prior research has primarily examined demographic and cultural biases, the ability of these models to accurately represent historical contexts remains largely underexplored. To address this gap, we introduce a benchmark for evaluating how TTI models depict historical contexts. The benchmark combines HistVis, a dataset of 30,000 synthetic images generated by three state-of-the-art diffusion models from carefully designed prompts covering universal human activities across multiple historical periods, with a reproducible evaluation protocol. We evaluate generated imagery across three key aspects: (1) Implicit Stylistic Associations: examining default visual styles associated with specific eras; (2) Historical Consistency: identifying anachronisms such as modern artifacts in pre-modern contexts; and (3) Demographic Representation: comparing generated racial and gender distributions against historically plausible baselines. Our findings reveal systematic inaccuracies in historically themed generated imagery, as TTI models frequently stereotype past eras by incorporating unstated stylistic cues, introduce anachronisms, and fail to reflect plausible demographic patterns. By providing a reproducible benchmark for historical representation in generated imagery, this work provides an initial step toward building more historically accurate TTI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。