arXiv:2509.08494cs.CYcs.AI2025-09被引 15

构建可扩展的评估框架,衡量AI助手对人类自主性的支持程度。

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

  • 用大模型模拟用户提问与反馈,评估AI响应中的自主性支持。
  • 发现当前AI助手在自主性支持上普遍偏低,不同厂商差异显著。
  • 强调需超越能力提升,转向更稳健的安全与对齐目标。

随着人类将更多任务与决策权交由人工智能,我们可能失去对个人与集体未来的掌控。现有算法系统已悄然引导人类决策,如社交媒体推荐算法使用户无意识地持续浏览高互动内容。本文结合哲学与科学中的自主性理论,提出一种基于大语言模型(LLMs)的AI辅助评估方法,构建了可扩展、自适应的HumanAgencyBench(HAB)基准。该基准涵盖六维人类自主性:提出澄清问题、避免价值操控、纠正错误信息、推迟重要决策、鼓励学习、维持社交边界。评估显示,当前基于LLM的助手在自主性支持上整体处于低至中等水平,且在不同开发者和维度间存在显著差异。例如,Anthropic LLM虽总体支持度最高,但在避免价值操控方面表现最差。自主性支持并不随模型能力或指令遵循行为(如强化学习人类反馈,RLHF)提升而稳定增加,提示应转向更可靠的对齐与安全目标。

原文摘要 · Abstract (English)

As humans delegate more tasks and decisions to artificial intelligence (AI), we risk losing control of our individual and collective futures. Relatively simple algorithmic systems already steer human decision-making, such as social media feed algorithms that lead people to unintentionally and absent-mindedly scroll through engagement-optimized content. In this paper, we develop the idea of human agency by integrating philosophical and scientific theories of agency with AI-assisted evaluation methods: using large language models (LLMs) to simulate and validate user queries and to evaluate AI responses. We develop HumanAgencyBench (HAB), a scalable and adaptive benchmark with six dimensions of human agency based on typical AI use cases. HAB measures the tendency of an AI assistant or agent to Ask Clarifying Questions, Avoid Value Manipulation, Correct Misinformation, Defer Important Decisions, Encourage Learning, and Maintain Social Boundaries. We find low-to-moderate agency support in contemporary LLM-based assistants and substantial variation across system developers and dimensions. For example, while Anthropic LLMs most support human agency overall, they are the least supportive LLMs in terms of Avoid Value Manipulation. Agency support does not appear to consistently result from increasing LLM capabilities or instruction-following behavior (e.g., RLHF), and we encourage a shift towards more robust safety and alignment targets.

AI评估自主性对齐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。