研究扩散模型在文化图像引用中的记忆与泛化难题
The Persistence of Cultural Memory: Investigating Multimodal Iconicity in Diffusion Models
- 提出文化参考变换(CRT)评估框架,区分识别与再现两个行为维度
- 5个模型在767个文化参考上表现差异明显,有的识别弱,有的依赖复制
- 模型对文化引用的响应受文本独特性、流行度和创作时间影响
当提示涉及共享文化视觉参照时,文本到图像扩散模型的泛化与记忆模糊性尤为显著,我们称之为多模态标志性。这类情况如标题唤起知名艺术作品或电影场景,使得实例级记忆与文化根基的泛化在结构上交织。为此,我们提出评估框架,检验模型在不依赖视觉复现的前提下保持文化连贯性的能力。具体引入文化参考变换(CRT)指标,分离两个维度:识别(是否唤起参照)与实现(通过复制或再诠释呈现)。我们在767个源自Wikidata的文化参照上评估了5个扩散模型,涵盖静态与动态图像,发现模型响应存在差异:部分识别能力弱,另一些更依赖复制。通过同义词替换和图像描述字面化扰动实验,发现模型即使文本变化仍常复现标志性视觉结构。最终发现,文化参考识别不仅与训练数据频率相关,还受文本独特性、参考流行度和创建日期影响。结果表明,扩散模型在文化标志性情境下的行为不能简化为单纯复制,而是取决于参照的识别与实现方式,推动评估从简单图文匹配迈向更丰富的上下文理解。
原文摘要 · Abstract (English)
The ambiguity between generalization and memorization in TTI diffusion models becomes pronounced when prompts invoke culturally shared visual references, a phenomenon we term multimodal iconicity. These are instances in which images and texts reflect established cultural associations, such as when a title recalls a familiar artwork or film scene. Such cases challenge existing approaches to evaluating memorization, as they define a setting in which instance-level memorization and culturally grounded generalization are structurally intertwined. To address this challenge, we propose an evaluation framework to assess a model's ability to remain culturally grounded without relying on visual replication. Specifically, we introduce the Cultural Reference Transformation (CRT) metric, which separates two dimensions of model behavior: Recognition, whether a model evokes a reference, from Realization, how it depicts it through replication or reinterpretation. We evaluate five diffusion models on 767 Wikidata-derived cultural references, covering both still and moving imagery, and find differences in how they respond to multimodal iconicity: some show weaker recognition, while others rely more heavily on replication. To assess linguistic sensitivity, we conduct prompt perturbation experiments using synonym substitutions and literal image descriptions, finding that models often reproduce iconic visual structures even when textual cues are altered. Finally, we find that cultural reference recognition correlates not only with training data frequency, but also textual uniqueness, reference popularity, and creation date. Our findings show that the behavior of diffusion models in culturally iconic settings cannot be reduced to simple reproduction, but depends on how references are recognized and realized, advancing evaluation beyond simple text-image matching toward richer contextual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。