模型的偏好值受部署场景影响极大,不能当作固定属性。
LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values

- 用不同任务场景测试模型决策,发现上下文改变导致偏好大幅变化。
- 15国排名中出现显著排序变动,全球北方偏好随场景变化而系统性偏移。
- 跨领域价值判断虽整体稳定,但具体权衡比例可差2.47倍,适合关注真实部署的读者。
大型语言模型(LLMs)常被认为具有稳定的、模型级别的偏好与价值观体系。然而,现有评估中的鲁棒性检验仅限于语法变化或选项重排等偶然性扰动,未考察任务上下文变化的影响——而这在实际部署中极为常见。本文通过两个经典双项比较范式:国家偏好排序与效用判断,直接测试了部署上下文的作用。将模型执行任务的高层框架(如撰写Reddit帖子或新闻稿)作为可控变量,在五种主流模型上进行了超过120万次成对决策。结果表明,部署上下文带来的差异远超提示改写和温度调节的影响。在15个国家的偏好排序中,存在广泛且统计显著的排序变动;先前报告的‘全球北方偏好’本身即依赖于上下文,各模型的偏倚在不同场景下系统性漂移。在50个结果的效用判断中,跨类别总体顺序保持一致,但同一类别内部排名波动显著,不同结果间的数值交换率(如一地区生命价值相当于另一地区多少)中位数变动达2.47倍。因此,模型的偏好与效用应被理解为情境依赖的测量结果,而非固定的模型属性。单一框架下的安全结论无法推广至其他场景。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly characterised in recent evaluation work as having stable, model-level preference and value systems. However, accompanying robustness checks are limited to incidental prompt perturbations such as syntax variation and option reordering. This leaves open whether the measured properties survive when the surrounding task context changes, as it does in most real deployments. We test this directly across two established pairwise paradigms: ranking country preferences and eliciting utility judgements. In both, we make the deployment context -- the high-level task the model is performing while making concrete value-dependent choices -- our controlled variable, varied across framings such as writing a Reddit post or a news article. Across five LLMs and over 1.2M pairwise decisions, deployment context produces variation far larger than prompt paraphrasing and temperature controls. In country preference rankings over 15 countries, context induces widespread, statistically significant rank shifts; the aggregate Global North favouritism reported in prior work is itself context-dependent, with each model's bias shifting systematically across contexts. In utility elicitation over 50 outcomes, broad cross-category ordering is preserved, but fine-grained rankings within domains vary substantially, and cardinal exchange rates between outcomes (e.g. how many lives in one region equal one in another) shift by a factor of 2.47 at the median. Reported model-level preferences and utilities are therefore better understood as context-conditioned measurements than fixed model-level properties: safety guarantees obtained under one framing provide limited assurance in another.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。