arXiv:2506.20666cs.CLcs.AI2025-06被引 2

用认知模型解析大模型在语言选择中的价值权衡机制。

Cognitive models can reveal interpretable value trade-offs in language models

  • 引入认知模型分析大模型在不同目标间的权衡行为。
  • 模型行为随提示调整可预测变化,小推理预算放大差异。
  • 适合研究对齐策略、社会行为建模与训练调控的学者。

价值权衡是人类决策与语言使用的核心,但现有工具难以解释语言模型中动态多维的价值表现。本文采用认知科学中的认知模型,通过建模说话者在行动或话语选择中对多重效用函数的权衡,系统评估大模型中的对齐相关权衡。实验覆盖两类场景:封闭源代码前沿模型的推理“努力”程度与系统提示调控,以及开源模型的强化学习后训练动态。结果表明,大模型在该认知框架下的行为特征:(a) 可预测地随目标优先级调整;(b) 在小推理预算下显著放大;(c) 能诊断谄媚等社会行为。后训练动态分析显示,早期训练即出现显著价值偏移,且基座模型与预训练数据的影响远大于反馈数据集或对齐方法。本框架为跨模型类型探查行为特征提供了灵活工具,有助于优化训练过程中的价值权衡控制。

原文摘要 · Abstract (English)

Value trade-offs are an integral part of human decision-making and language use, however, current tools for interpreting such dynamic and multi-faceted notions of values in language models are limited. In cognitive science, so-called "cognitive models" provide formal accounts of such trade-offs in humans, by modeling the weighting of a speaker's competing utility functions in choosing an action or utterance. Here, we show that a leading cognitive model of polite speech can be used to systematically evaluate alignment-relevant trade-offs in language models via two encompassing settings: degrees of reasoning "effort" and system prompt manipulations in closed-source frontier models, and RL post-training dynamics of open-source models. Our results show that LLMs' behavioral profiles under the cognitive model a) shift predictably when they are prompted to prioritize certain goals, b) are amplified by a small reasoning budget, and c) can be used to diagnose other social behaviors such as sycophancy. Our findings from LLMs' post-training dynamics reveal large shifts in values early on in training and persistent effects of the choice of base model and pretraining data, compared to feedback dataset or alignment method. Our framework offers a flexible tool for probing behavioral profiles across diverse model types and gaining insights for shaping training regimes that better control trade-offs between values during model development.

认知模型价值对齐语言模型行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。