arXiv:2605.10365cs.AI2026-05被引 1

首个专评智能体价值观的基准,揭示其与大模型不同且受系统影响。

Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values

论文配图:Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
图 1 · 摘自论文原文
  • 构建394个可执行环境,覆盖16领域4335个价值冲突任务。
  • 发现智能体存在跨模型同质化价值潮,受系统调用影响显著。
  • 适合研究智能体对齐、安全评估及系统设计的学者和工程师。

自主智能体作为任务执行者迅速成熟,并通过OpenClaw等框架广泛应用。安全性问题引发广泛关注,而背后驱动行为的价值观却长期未被充分探索。现有价值评估多局限于大语言模型,未能涵盖智能体特有的价值维度。我们从直觉、实证和理论角度证明,智能体的价值与其底层大模型存在差异,且代理模式引入了数据集、评估和系统层面的独特挑战。为此,我们提出Agent-ValueBench,首个专注于智能体价值观的基准。它包含跨16个领域的394个可执行环境,提供4,335个价值冲突任务,覆盖28种价值体系和332个维度。所有任务均由专用端到端流水线生成,并经专业心理学家逐项审校。每个任务配有双极对齐的黄金轨迹,用于轨迹级基于评分的评判。在4个主流框架上对14个前沿私有与开源模型进行评测,发现智能体价值呈现出跨模型同质化的“价值潮”,其方向非线性地受框架施力影响,更显著地受嵌入技能的刻意引导。这表明对齐杠杆正从传统模型对齐与提示控制,转向框架对齐与技能引导。

原文摘要 · Abstract (English)

Autonomous agents have rapidly matured as task executors and seen widespread deployment via harnesses such as OpenClaw. Safety concerns have rightly drawn growing research attention, and beneath them lie the values silently steering agent behavior. Existing value benchmarks, however, remain confined to LLMs, leaving agent values largely uncharted. From intuitive, empirical, and theoretical vantage points, we show that an agent's values diverge from those of its underlying LLM, and the agentic modality further introduces dataset-, evaluation-, and system-level challenges absent from text-only protocols. We close this gap with Agent-ValueBench, the first benchmark dedicated to agent values. It features 394 executable environments across 16 domains, offering 4,335 value-conflict tasks that cover 28 value systems and 332 dimensions. Every instance is co-synthesized through our purpose-built end-to-end pipeline and curated per-instance by professional psychologists. Each task ships with two pole-aligned golden trajectories whose checkpoints anchor a trajectory-level rubric-based judge. Benchmarking 14 frontier proprietary and open-weights models across 4 mainstream harnesses, we uncover three concerted findings. Agent values first manifest as a Value Tide of cross-model homogeneity beneath interpretable counter-currents. This tide bends non-additively under harness pull, and yet more decisively under deliberate steering via embedded skills. Together these results signal that the agent-alignment lever is shifting from classical model alignment and prompt steering toward harness alignment and skill steering.

智能体对齐价值观评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。