用情境题探测并操控大模型的文化倾向,揭示隐性价值观耦合现象。
Scenario-based Probing and Steering Cultural Values in Large Language Models--Extended Version
- 将文化问题转为行为困境场景,通过词元概率捕捉隐性偏好。
- 在3个模型、4种文化中验证可调节性,发现维度间存在耦合效应。
- 适合研究模型文化偏见或需跨文化对齐的AI系统开发者。
大型语言模型在跨文化场景中部署,却常反映训练数据带来的同质化价值取向。现有评估多依赖直接提问,易得中立或安全回应,难以捕捉深层偏好。本文提出基于世界价值观调查(WVS)双轴框架的探测与调控方法,将社会价值问题转化为情境化行为困境,通过提取词元级概率来度量隐含价值观,并应用激活调控,辅以国家条件提示,实现无需重训练的行为调整。在三个开源模型和四种目标文化中,发现显著的可调性差异及潜在耦合现象:一个维度的干预会引发另一维度的变化。这种耦合与人类WVS数据中的相关性一致,且在激活、提示及混合调控中均存在,限制了轴向独立对齐,但通用任务性能基本保持。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are deployed across cultural contexts but often reflect homogenized values inherited from training data. Evaluations of cultural alignment typically rely on direct prompting with survey-style questions, which frequently elicit neutral or safety-aligned responses and fail to capture underlying model preferences. We propose a framework for probing and steering latent cultural representations in LLMs along the two Inglehart--Welzel axes of the World Values Survey (WVS). By translating social value questions into scenario-based behavioral dilemmas, we extract token-level probabilities to measure implicit values and apply activation steering, optionally combined with country-conditioned prompting, to shift model behavior without retraining. Across three open-source LLMs and four target cultures, we find substantial variation in steerability and identify latent entanglement, where interventions along one cultural dimension induce shifts along another. This coupling mirrors correlations in human WVS data and persists across activation, prompt, and hybrid steering. It constrains axis-independent alignment, though general task performance is largely preserved.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。