arXiv:2502.05331cs.CL2025-02NAACL被引 2

用微调的LLM当时间胶囊,追踪书籍中的社会偏见演变

Fine-Tuned LLMs are "Time Capsules" for Tracking Societal Bias Through Books

  • 用各年代小说微调LLM,通过提示词分析偏见变化
  • 女性领导力占比从8%升至22%,同性关系提及率从0%到10%
  • 揭示偏见源于文本内容,适合人文与AI交叉研究者

书籍虽富含文化洞察,但也映射其时代的社会偏见,这些偏见可能被大语言模型(LLMs)在训练中习得并延续。我们提出一种新方法,利用微调的LLM追踪和量化这些偏见。构建了包含593部虚构小说的BookPAGE语料库,覆盖1950至2019年七个十年,通过为每十年的书籍微调LLM,并使用定向提示词,考察性别、性取向、种族和宗教相关偏见的变化。结果显示,基于特定年代书籍训练的模型展现出与其时代相符的偏见特征,存在渐进趋势和显著转折。例如,女性担任领导角色的描述比例从1950年代的8%上升至2010年代的22%,1990年代出现明显跃升(4%至12%),可能与第三波女权主义相关;同性关系提及率从1980年代的0%增至2000年代的10%,反映LGBTQ+可见度提升;令人担忧的是,伊斯兰教负面刻画在2000年代由26%升至38%,或与9/11后情绪有关。重要的是,我们证明这些偏见主要源自书籍内容,而非模型架构或初始训练。本研究为社会偏见演化提供了新视角,融合了人工智能、文学研究与社会科学。

原文摘要 · Abstract (English)

Books, while often rich in cultural insights, can also mirror societal biases of their eras - biases that Large Language Models (LLMs) may learn and perpetuate during training. We introduce a novel method to trace and quantify these biases using fine-tuned LLMs. We develop BookPAGE, a corpus comprising 593 fictional books across seven decades (1950-2019), to track bias evolution. By fine-tuning LLMs on books from each decade and using targeted prompts, we examine shifts in biases related to gender, sexual orientation, race, and religion. Our findings indicate that LLMs trained on decade-specific books manifest biases reflective of their times, with both gradual trends and notable shifts. For example, model responses showed a progressive increase in the portrayal of women in leadership roles (from 8% to 22%) from the 1950s to 2010s, with a significant uptick in the 1990s (from 4% to 12%), possibly aligning with third-wave feminism. Same-sex relationship references increased markedly from the 1980s to 2000s (from 0% to 10%), mirroring growing LGBTQ+ visibility. Concerningly, negative portrayals of Islam rose sharply in the 2000s (26% to 38%), likely reflecting post-9/11 sentiments. Importantly, we demonstrate that these biases stem mainly from the books' content and not the models' architecture or initial training. Our study offers a new perspective on societal bias trends by bridging AI, literary studies, and social science research.

社会偏见LLM文本分析跨学科

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。