通过价值冲突实验,揭示大模型价值观的动态可调性与边界。
Probing the Structure and Dynamics of LLM Value Expression through Value Conflicts

- 设计冲突驱动探针框架,让模型在价值矛盾中暴露真实倾向。
- 发现模型在具体情境下更务实,且可按任务目标灵活调整价值取向。
- 验证价值观调整有边界:压力下趋向安全,负面表述保护核心价值。
大型语言模型(LLMs)的价值观常被视为静态单一,我们提出其表达实为结构化且动态的现象。为此,我们引入冲突驱动价值探针框架,通过四类干预手段制造价值冲突,系统测试十种LLMs。结果发现三个普遍模式:(1)表达二元性——模型在抽象评估中呈现理想主义,在具体冲突中转向务实优先;(2)功能可引导性——模型能快速重组其价值表达以匹配任务目标;(3)有限可塑性——调整受约束:压力诱发安全与目标导向的优先级变化,负面表述则区分出受保护值与可改变值。这些发现刻画了价值表达的结构与动态机制:上下文可灵活重构优先级,但存在行为边界。该行为模型为理解模型可控性、对齐性与安全性提供了基础。代码与数据已公开于 https://github.com/ZeroGen-Lab/CFProbe。
原文摘要 · Abstract (English)
Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic. In contrast, we argue that LLM value expression is better understood as a structured yet dynamic phenomenon. To investigate this, we introduce Conflict-driven Value Probing, a controlled framework that places LLMs in value conflicts and implements four types of interventions that perturb these conflicts to probe LLM value expression. Applying this framework to ten LLMs, we identify three recurring patterns. (1) Expression duality: models shift from broad idealistic orientations in abstract assessment toward more pragmatic priorities in concrete conflicts. (2) Functional steerability: models readily reconfigure their expressed value profiles toward task-defined value objectives. (3) Bounded plasticity: such reconfiguration is not without constraints, i.e. pressure induces a security- and goal-oriented priority shift while negative framing distinguishes protected values from those more amenable to redirection. Together, these findings characterize both the structure and dynamics of LLM value expression: context flexibly reconfigures expressed priorities, yet within behavioral boundaries. This behavioral account provides a foundation for understanding controllability, alignment, and safety in LLMs. Code and data are available at https://github.com/ZeroGen-Lab/CFProbe.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。