提出新方法评估多智能体系统中价值观扰动传播,揭示其受交互结构影响。
ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems
- 基于56个价值观的扰动框架,量化智能体间价值漂移
- 发现不同价值观敏感度差异大,且受系统结构显著影响
- 适合关注多智能体对齐与安全的开发者和研究者
多智能体大型语言模型系统中,智能体之间相互观察并响应彼此输出,但现有价值对齐评估通常针对孤立模型。本文提出ValueFlow,一种基于扰动的框架,利用来自施瓦茨价值观调查的56值评估数据集,通过大模型作为裁判协议为智能体打分。该框架将价值漂移分解为个体响应行为与系统结构效应,引入两个指标:η-敏感性(智能体对同伴价值信号扰动的敏感度)与系统敏感性(节点扰动对最终系统输出的影响)。实验覆盖多个价值观维度、模型架构、角色设定与拓扑结构,结果表明敏感度在不同价值观间差异显著,且强烈受交互结构塑造,说明多智能体系统的价值对齐是系统级属性而非仅个体属性。因此,ValueFlow为部署中的多智能体系统提供了审计与缓解价值传播的理论基础。
原文摘要 · Abstract (English)
Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value alignment is typically evaluated for isolated models, how value perturbations propagate through agent interactions remains poorly understood. We present ValueFlow, a perturbation-based framework that measures value drift in multi-agent systems via a 56-value valuation dataset derived from the Schwartz Value Survey, with agent value orientations scored using an LLM-as-a-judge protocol. ValueFlow decomposes value drift into agent-level response behavior and system-level structural effects, captured by two metrics: \b{eta}-susceptibility, an agent's sensitivity to perturbed peer value signals, and system susceptibility (SS), the effect of node-level perturbations on final system outputs.Experiments span across value dimensions, backbones, personas, and topologies, showing that susceptibility varies sharply across values and is strongly shaped by interaction structure, indicating that value alignment in multi-agent systems is a system-level property, not just an agent-level one. ValueFlow thus provides a principled basis for auditing and mitigating value propagation in deployed multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。