用问卷微调让大模型更符合人类价值观,效果显著。
Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
- 通过问卷微调让模型回答更贴近人类价值观。
- 微调后模型在问卷和情境任务中行为均明显改变。
- 适合关注价值观对齐与模型行为控制的研究者。
大型语言模型隐含人类价值观偏好,但通常需大量数据进行引导。本文探究一种简单方法:能否通过让模型针对价值观问卷正确作答,来改变其下游行为中的价值取向?我们首先构建多个开源大模型在20种人类价值观上的价值画像作为基线。随后研究通过价值观问卷微调是否能改变模型的价值体系。评估方式包括:1)在保留的问卷问题上观察答案变化;2)在非领域情境(如基于Reddit帖子构建的道德判断数据集、文本冒险游戏)中评估行为变化。结果表明,该方法不仅能有效调整模型在问卷中的回答,还能在隐式下游任务中产生显著的价值对齐效应。
原文摘要 · Abstract (English)
Large language models implicitly encode preferences over human values, yet steering them often requires large training data. In this work, we investigate a simple approach: Can we reliably modify a model's value system in downstream behavior by training it to answer value survey questions accordingly? We first construct value profiles of several open-source LLMs by asking them to rate a series of value-related descriptions spanning 20 distinct human values, which we use as a baseline for subsequent experiments. We then investigate whether the value system of a model can be governed by fine-tuning on the value surveys. We evaluate the effect of finetuning on the model's behavior in two ways; first, we assess how answers change on in-domain, held-out survey questions. Second, we evaluate whether the model's behavior changes in out-of-domain settings (situational scenarios). To this end, we construct a contextualized moral judgment dataset based on Reddit posts and evaluate changes in the model's behavior in text-based adventure games. We demonstrate that our simple approach can not only change the model's answers to in-domain survey questions, but also produces substantial shifts (value alignment) in implicit downstream task behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。