arXiv:2510.26707cs.CLcs.CY2025-10Transactions of th…被引 11

追踪大模型后训练中价值观演变过程,发现微调阶段决定价值取向

Value Drifts: Tracing Value Alignment During LLM Post-Training

  • 通过分解训练算法与数据影响,量化价值漂移的时机与幅度
  • 微调阶段确立模型价值观,后续偏好优化几乎不改变已形成的价值
  • 不同偏好优化算法即使在相同数据下也导致不同价值结果,适合对齐研究者参考

随着大语言模型在社会中作用日益重要,它们不仅需要运用通用知识,还需与特定人类价值体系对齐。因此,研究模型与人类价值观的对齐成为关键课题。然而,以往工作多聚焦于完全训练后模型的对齐评估,忽视了模型学习表达人类价值观的过程动态。本文研究大模型后训练过程中价值对齐的产生机制与阶段。分析解耦了后训练算法与数据集的影响,测量了训练期间价值漂移的幅度与时序。在Llama-3和Qwen-3不同规模模型上,使用主流监督微调(SFT)和偏好优化数据集及算法进行实验,发现SFT阶段通常奠定模型价值观,后续偏好优化很少重新对齐这些价值。进一步使用可控制价值的合成偏好数据集,发现即使偏好数据保持不变,不同偏好优化算法也会导致不同的价值对齐结果。研究为理解后训练中价值观学习机制提供实证洞察,有助于数据筛选、模型与算法选择以提升模型与人类价值观的对齐效果。

原文摘要 · Abstract (English)

As LLMs occupy an increasingly important role in society, they are more and more confronted with questions that require them not only to draw on their general knowledge but also to align with certain human value systems. Therefore, studying the alignment of LLMs with human values has become a crucial field of inquiry. Prior work, however, mostly focuses on evaluating the alignment of fully trained models, overlooking the training dynamics by which models learn to express human values. In this work, we investigate how and at which stage value alignment arises during the course of a model's post-training. Our analysis disentangles the effects of post-training algorithms and datasets, measuring both the magnitude and time of value drifts during training. Experimenting with Llama-3 and Qwen-3 models of different sizes and popular supervised fine-tuning (SFT) and preference optimization datasets and algorithms, we find that the SFT phase generally establishes a model's values, and subsequent preference optimization rarely re-aligns these values. Furthermore, using a synthetic preference dataset that enables controlled manipulation of values, we find that different preference optimization algorithms lead to different value alignment outcomes, even when preference data is held constant. Our findings provide actionable insights into how values are learned during post-training and help to inform data curation, as well as the selection of models and algorithms for preference optimization to improve model alignment to human values.

价值对齐后训练大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。