揭示大模型文化生成偏差源于预训练数据记忆,而非真正理解。
Attributing Culture-Conditioned Generations to Pretraining Corpora
- 通过分析预训练数据中文化-实体关联模式,识别生成偏差来源。
- 高频文化有更多记忆性生成,低频文化几乎无生成内容。
- 模型偏好频繁出现的词汇,无视文化相关性,存在固有偏见。
在开放生成任务如叙事写作或对话中,大型语言模型常表现出文化偏见,对不常见文化知识有限且生成模板化内容。近期研究指出这些偏见可能源于预训练语料库中文化代表性不均。本文通过分析模型如何根据预训练数据模式将实体与文化关联,探究预训练导致文化条件生成偏见的机制。提出MEMOed框架(基于预训练文档的记忆化检测),用于判断某一文化生成是否源于记忆。在110个文化的饮食与服饰生成上应用MEMOed发现:预训练数据中高频文化产生更多记忆符号生成,而部分低频文化则无生成;同时,模型无论条件文化如何,均倾向生成频率极高的实体,反映对高频预训练词的固有偏好。期望本框架与发现能推动对模型性能与预训练数据关联性的进一步研究。
原文摘要 · Abstract (English)
In open-ended generative tasks like narrative writing or dialogue, large language models often exhibit cultural biases, showing limited knowledge and generating templated outputs for less prevalent cultures. Recent works show that these biases may stem from uneven cultural representation in pretraining corpora. This work investigates how pretraining leads to biased culture-conditioned generations by analyzing how models associate entities with cultures based on pretraining data patterns. We propose the MEMOed framework (MEMOrization from pretraining document) to determine whether a generation for a culture arises from memorization. Using MEMOed on culture-conditioned generations about food and clothing for 110 cultures, we find that high-frequency cultures in pretraining data yield more generations with memorized symbols, while some low-frequency cultures produce none. Additionally, the model favors generating entities with extraordinarily high frequency regardless of the conditioned culture, reflecting biases toward frequent pretraining terms irrespective of relevance. We hope that the MEMOed framework and our insights will inspire more works on attributing model performance on pretraining data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。