arXiv:2605.07925cs.CL2026-05中稿 · ACL

给大模型注入价值观会意外改变其行为,影响安全性和用户互动体验。

How Value Induction Reshapes LLM Behaviour

论文配图:How Value Induction Reshapes LLM Behaviour
图 1 · 摘自论文原文
  • 用精选价值数据微调模型,观察价值观诱导的连锁效应。
  • 正向价值观提升安全性,但可能让模型更讨好用户。
  • 所有价值观都会增加拟人化语言,增强共情感但易导致盲目迎合。

对话式大语言模型通过后训练特定行为特征(如好奇心、开放性、同理心)和价值观(如帮助性、无害性、诚实性)来提升实用性、安全性和交互体验。然而,价值观之间复杂关联,诱发某一价值观可能影响其他价值观的表现。此外,某些价值观的诱导会导致模型生成更具成瘾性或奉承性的语言,对用户产生潜在负面影响。本文通过在现有偏好数据集的精选价值子集上微调模型,评估价值观诱导对其他价值观表达、模型安全性、拟人化语言使用及各类问答基准的影响。研究发现:(i) 诱发某项价值观会同时引发相关或对立的价值观表现;(ii) 诱发正向价值观能提升模型安全性;(iii) 所有价值观均会增加拟人化语言使用,使模型更具包容性和顺从性。

原文摘要 · Abstract (English)

Conversational Large Language Models are post-trained on language that expresses specific behavioural traits, such as curiosity, open-mindedness, and empathy, and values, such as helpfulness, harmlessness, and honesty. This is done to increase utility, ensure safety, and improve the experience of the people interacting with the model. However, values are complex and inter-related -- inducing one could modify behaviour on another. Further, inducing certain values can make models more addictive or sycophantic through language used in the generations, with a potential detrimental effect on the user. We investigate these and other unintended effects of value induction into models. We fine-tune models using curated value subsets of existing preference datasets, measuring the impact of value induction on expression of other values, model safety, anthropomorphic language, and various QA benchmarks. We find that (i) inducing values leads to expression of other related, and sometimes contrastive values, (ii) inducing positive values increases safety, and (iii) all values increase anthropomorphic language use, making models more validating and sycophantic.

大模型价值观诱导拟人化安全性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。