安全训练让大模型压制文化知识,且不平等。
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
- 发现模型在拒绝时内部输出仍保留文化认知,称'对齐否决'
- 安全训练导致文化知识被压制,非彻底删除
- 适合关注伦理、公平与模型对齐的读者
在16个中东和北非(MENA)国家、26个模型及153万份人类调查数据下,我们发现当对齐训练与文化价值观冲突时,模型并非消除文化知识,而是抑制其表达:在拒绝时刻,模型内部的逻辑分布与人类调查数据相关性高于自由生成答案。这种现象称为‘对齐否决’。我们区分了‘抑制失败’(内部准确但输出被阻断)与‘表征偏差失败’(编码本身偏离人类价值观),并证明二者需不同干预策略。该机制存在不公:安全代价高达37.6%,最佳与最差服务国家间对齐质量差距达19.8%;使用母语提示反而扩大差距。稀疏自编码器分析结合对齐阶段对比,识别出Tulu-3-8B中可能介导抑制的候选特征位于DPO阶段。该机制真实存在,其代价不均,而决定保护什么并非技术问题。
原文摘要 · Abstract (English)
What happens inside a language model when alignment training conflicts with a cultural value it encodes? Across 16 MENA countries, 26 models, and 1.53M human survey responses, we show the answer is suppression, not erasure: at the moment of refusal, a model's internal logit distribution correlates with human survey data more strongly than its freely generated answers. We call this the alignment veto. We distinguish suppression failures (accurate internal distributions blocked at output) from representational bias failures (the encoding itself diverges from human values), and show the two require different interventions. The gate is inequitable: the safety tax reaches 37.6%, with a 19.8% alignment-quality gap between best- and worst-served nations, and native-language prompting widens rather than closes it. Sparse autoencoder analysis corroborated by comparisons across alignment stages identifies a candidate DPO-stage feature mediating suppression in Tulu-3-8B. The gate is real, its costs are unequal, and deciding what it should protect is not a technical question.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。