arXiv:2605.26365cs.CL2026-05被引 1

通过行为探测让大模型暴露隐含文化偏好并精准调整。

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

论文配图:Cultural Value Alignment Via Latent Activation Steering in Large Language Models
图 1 · 摘自论文原文
  • 用300个情境难题提取隐层概率,揭示模型真实文化立场。
  • 无需训练即可在推理时调整文化倾向,实现动态干预。
  • 发现文化维度相互纠缠,精准对齐存在根本限制。

大型语言模型常表现出同质化的文化视角。尽管世界价值观调查(WVS)提供了人类价值观的黄金标准,但传统直接提示LLM回答WVS问题往往无法触及模型潜在的文化深度,导致安全对齐下的拒绝或中立回应。本文提出一种通用的文化评估与干预框架,从抽象提问转向基于情境的行为探测。通过提取300个情境困境中的隐层标记概率,绕过表层对齐,映射出模型潜在的文化价值坐标。进一步引入激活调节技术,在前向传播过程中调整内部对齐,无需重训。在多个大模型上,我们发现其适应性存在显著差异,并揭示出一种普遍存在的隐层纠缠现象:在某一文化维度上的干预会引发其他维度的同步偏移。结果表明,文化价值观以耦合结构编码,限制了精确对齐的可能性。本工作建立了一个计算高效的模型文化调节框架,凸显了在全球价值观导航中模型内在结构的复杂性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping human values, traditional direct prompting of LLMs on WVS often fails to access the model's latent cultural depth, leading to safety-aligned refusals or neutral responses. Here, we propose a generalizable framework for cultural evaluation and intervention that transitions from abstract queries to scenario-based behavioral probing. By extracting implicit token probabilities across 300 situational dilemmas, we bypass surface-level alignment to map the latent coordinates of LLMs cultural value. We further introduce activation steering to shift these internal alignments during the forward pass without retraining. Across multiple LLMs, we find substantial variation in adaptability and uncover a consistent phenomenon of latent entanglement, where interventions along one cultural dimension induce shifts along another. These results suggest that cultural values are encoded as coupled structures, limiting precise alignment. This work establishes a computationally efficient framework for cultural steering, highlighting the structural complexities when navigating global value with LLMs.

文化对齐激活调节价值观建模大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。