发现大模型隐含价值结构与人类不同,提出轻量级对齐新方法。
Are the Values of LLMs Structurally Aligned with Humans? A Causal Perspective
- 构建大模型的潜在因果价值图,揭示其内在价值结构。
- 在Gemma-2B-IT和Llama3-8B-IT上验证方法有效,可精准控制输出。
- 相比传统方法更轻量、可解释,适合实际部署场景。
随着大语言模型(LLMs)在关键应用中日益普及,使其行为与人类价值观对齐面临重大挑战。现有方法如基于人类反馈的强化学习(RLHF)通常仅关注有限的粗粒度价值观,且资源消耗大。此外,这些价值观之间的关联关系不明确,导致价值调控结果难以解释。本文认为,大模型的价值维度背后存在一个潜在的因果价值图,尽管经过对齐训练,该结构仍与人类价值体系显著不同。我们利用这一因果价值图,指导两种轻量级价值调控方法:基于角色的提示(role-based prompting)和稀疏自编码器(SAE)调控,有效缓解了意外副作用。此外,SAE提供了更细粒度的价值调控能力。在Gemma-2B-IT和Llama3-8B-IT上的实验表明,所提方法具有良好的有效性与可控性。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly integrated into critical applications, aligning their behavior with human values presents significant challenges. Current methods, such as Reinforcement Learning from Human Feedback (RLHF), typically focus on a limited set of coarse-grained values and are resource-intensive. Moreover, the correlations between these values remain implicit, leading to unclear explanations for value-steering outcomes. Our work argues that a latent causal value graph underlies the value dimensions of LLMs and that, despite alignment training, this structure remains significantly different from human value systems. We leverage these causal value graphs to guide two lightweight value-steering methods: role-based prompting and sparse autoencoder (SAE) steering, effectively mitigating unexpected side effects. Furthermore, SAE provides a more fine-grained approach to value steering. Experiments on Gemma-2B-IT and Llama3-8B-IT demonstrate the effectiveness and controllability of our methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。