揭示语言模型架构如何影响偏见传播,提出系统性分析方法
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
- 通过对比n-gram与Transformer模型,分析架构对偏见传播的影响
- 发现n-gram对上下文窗口敏感,而Transformer更具鲁棒性
- 指出训练数据时间来源和特定偏见类型会显著放大偏差
当前关于语言模型偏见的研究多聚焦于数据质量,较少关注模型架构及数据的时间动态影响。更关键的是,很少有研究系统探究偏见的根源。本文基于比较行为理论,提出一种方法来解析训练数据与模型架构在语言建模中偏见传播的复杂交互。结合近期将Transformer与n-gram语言模型关联的研究,我们评估了数据、模型设计选择与时间动态对偏见传播的影响。结果表明:(1) n-gram模型在偏见传播上对上下文窗口大小高度敏感,而Transformer表现出架构鲁棒性;(2) 训练数据的时间来源显著影响偏见程度;(3) 不同模型架构对受控偏见注入反应不同,某些偏见(如性取向)被显著放大。随着语言模型广泛应用,本研究强调需从数据与模型双重维度追溯偏见根源,而非仅关注表象,以减少潜在伤害。
原文摘要 · Abstract (English)
Current research on bias in language models (LMs) predominantly focuses on data quality, with significantly less attention paid to model architecture and temporal influences of data. Even more critically, few studies systematically investigate the origins of bias. We propose a methodology grounded in comparative behavioral theory to interpret the complex interaction between training data and model architecture in bias propagation during language modeling. Building on recent work that relates transformers to n-gram LMs, we evaluate how data, model design choices, and temporal dynamics affect bias propagation. Our findings reveal that: (1) n-gram LMs are highly sensitive to context window size in bias propagation, while transformers demonstrate architectural robustness; (2) the temporal provenance of training data significantly affects bias; and (3) different model architectures respond differentially to controlled bias injection, with certain biases (e.g. sexual orientation) being disproportionately amplified. As language models become ubiquitous, our findings highlight the need for a holistic approach -- tracing bias to its origins across both data and model dimensions, not just symptoms, to mitigate harm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。