arXiv:2508.11414cs.CL2025-08综述

用问卷微调让大模型更符合人类价值观,效果显著。

Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions

  • 通过问卷微调让模型回答更贴近人类价值观。
  • 微调后模型在问卷和情境任务中行为均明显改变。
  • 适合关注价值观对齐与模型行为控制的研究者。

大型语言模型隐含人类价值观偏好,但通常需大量数据进行引导。本文探究一种简单方法:能否通过让模型针对价值观问卷正确作答,来改变其下游行为中的价值取向?我们首先构建多个开源大模型在20种人类价值观上的价值画像作为基线。随后研究通过价值观问卷微调是否能改变模型的价值体系。评估方式包括:1)在保留的问卷问题上观察答案变化;2)在非领域情境(如基于Reddit帖子构建的道德判断数据集、文本冒险游戏)中评估行为变化。结果表明,该方法不仅能有效调整模型在问卷中的回答,还能在隐式下游任务中产生显著的价值对齐效应。

原文摘要 · Abstract (English)

Large language models implicitly encode preferences over human values, yet steering them often requires large training data. In this work, we investigate a simple approach: Can we reliably modify a model's value system in downstream behavior by training it to answer value survey questions accordingly? We first construct value profiles of several open-source LLMs by asking them to rate a series of value-related descriptions spanning 20 distinct human values, which we use as a baseline for subsequent experiments. We then investigate whether the value system of a model can be governed by fine-tuning on the value surveys. We evaluate the effect of finetuning on the model's behavior in two ways; first, we assess how answers change on in-domain, held-out survey questions. Second, we evaluate whether the model's behavior changes in out-of-domain settings (situational scenarios). To this end, we construct a contextualized moral judgment dataset based on Reddit posts and evaluate changes in the model's behavior in text-based adventure games. We demonstrate that our simple approach can not only change the model's answers to in-domain survey questions, but also produces substantial shifts (value alignment) in implicit downstream task behavior.

价值观对齐微调行为控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。