arXiv:2603.23301cs.CL2026-03被引 4

用可解释性技术让大模型生成更符合特定文化的回答。

Steering LLMs for Culturally Localized Generation

  • 通过稀疏自编码器识别文化相关特征,构建可调控的文化嵌入向量。
  • 相比传统提示法,能显著激发罕见的长尾文化概念,提升文化契合度。
  • 适合研究模型偏见或需要跨文化生成的开发者与研究人员。

大语言模型在全球部署,但其输出常偏向训练数据丰富的文化。现有文化本地化方法如提示工程或后训练对齐均为黑箱,难以控制,且无法判断失败是知识缺失还是表达诱导不当。本文采用机制可解释性技术,揭示并操控模型中的文化表征。利用稀疏自编码器,我们识别出编码文化关键信息的可解释特征,并聚合为文化嵌入(CuE)。CuE既可用于分析未明确提示下的隐含文化偏见,也可用于构建白盒引导干预。在多个模型上验证,基于CuE的引导显著提升了文化忠实度,并激发了远比提示法更多样、更稀有的长尾文化概念。值得注意的是,该方法与黑箱本地化方法互补,在提示增强输入基础上仍能带来增益,表明模型确实受益于更好的诱导策略,而非普遍缺乏长尾知识表示,但效果因文化而异。结果为理解大模型中的文化表征提供了诊断视角,也提供了一种可控的文化引导方法。

原文摘要 · Abstract (English)

LLMs are deployed globally, yet produce responses biased towards cultures with abundant training data. Existing cultural localization approaches such as prompting or post-training alignment are black-box, hard to control, and do not reveal whether failures reflect missing knowledge or poor elicitation. In this paper, we address these gaps using mechanistic interpretability to uncover and manipulate cultural representations in LLMs. Leveraging sparse autoencoders, we identify interpretable features that encode culturally salient information and aggregate them into Cultural Embeddings (CuE). We use CuE both to analyze implicit cultural biases under underspecified prompts and to construct white-box steering interventions. Across multiple models, we show that CuE-based steering increases cultural faithfulness and elicits significantly rarer, long-tail cultural concepts than prompting alone. Notably, CuE-based steering is complementary to black-box localization methods, offering gains when applied on top of prompt-augmented inputs. This also suggests that models do benefit from better elicitation strategies, and don't necessarily lack long-tail knowledge representation, though this varies across cultures. Our results provide both diagnostic insight into cultural representations in LLMs and a controllable method to steer towards desired cultures.

文化生成可解释性大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。