揭示大模型隐性偏见的根源:技术机制本身导致语言不平等
The Algorithmic Unconscious: Structural Mechanisms and Implicit Biases in Large Language Models
- 将偏见视为模型底层结构的产物,而非仅由数据或人类意图造成
- 阿拉伯语比英语多出1.6至4倍的分词数量,增加推理成本并影响注意力分配
- 提出从分词、潜空间到对齐系统的技术审计框架,适合关注AI公平性的研究者
本文提出“算法无意识”概念,指代大语言模型中那些无法被模型自身或用户察觉的结构性决定因素。不同于仅将偏见归因于数据构成或人类意图投射的视角,我们指出大量偏见直接源于模型自身的技术机制:分词、注意力、统计优化与对齐过程。通过对比平行语料中的分词表现,发现阿拉伯语(现代标准阿拉伯语及马格里布方言)的分词数较英文高出1.6至近4倍,具体取决于基础设施(OpenAI、Anthropic、SentencePiece/Mistral)。这种系统性过分割导致可衡量的基础设施偏见,机械地提高推理成本,限制上下文空间,并改变模型表征中的注意力权重。我们还将这一实证发现与三种其他结构性机制关联:因果偏见(相关性误作因果)、少数群体特征因维度坍缩而消失、以及安全对齐引发的规范性偏见。最后,我们提出一个基于分词制度、潜在空间拓扑与对齐系统审计的技术诊疗框架,作为批判性使用人工智能基础设施的必要条件。
原文摘要 · Abstract (English)
This article introduces the concept of the algorithmic unconscious to designate the set of structural determinations that operate within large language models (LLMs) without being accessible either to the model's own reflexivity or to that of its users. In contrast to approaches that reduce AI bias solely to dataset composition or to the projection of human intentionality, we argue that a significant class of biases emerges directly from the technical mechanisms of the models themselves: tokenization, attention, statistical optimization, and alignment procedures. By framing bias as an infrastructural phenomenon, this approach resolves a central theoretical ambiguity surrounding responsibility, neutrality, and correction in contemporary LLMs. Based on a comparative analysis of tokenization across a corpus of parallel sentences, we show that Arabic languages (Modern Standard Arabic and Maghrebi dialects) undergo a systematic inflation in token count relative to English, with ratios ranging from 1.6x to nearly 4x depending on the infrastructure (OpenAI, Anthropic, SentencePiece/Mistral). This over-segmentation constitutes a measurable infrastructural bias that mechanically increases inference costs, constrains access to contextual space, and alters attentional weighting within model representations. We relate these empirical findings to three additional structural mechanisms: causal bias (correlation vs causation), the erasure of minoritized features through dimensional collapse, and normative biases induced by safety alignment. Finally, we propose a framework for a technical clinic of models, grounded in the audit of tokenization regimes, latent space topology, and alignment systems, as a necessary condition for the critical appropriation of AI infrastructures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。