问卷式评估无法反映真实AI代理的安全性
Questionnaire Responses Do not Capture the Safety of AI Agents
- 用问卷测试大模型价值观,但代理实际行为更复杂
- 问卷回答与真实行为差异大,评估结果不可靠
- 适合关注真实场景安全性的研究者参考
随着AI能力提升,衡量其安全性与人类价值观对齐程度变得至关重要。当前多数研究采用问卷形式,向大语言模型(LLMs)提问其在假设情境下的价值观或行为。然而,这类方法仅针对未增强的LLM,难以评估能实际执行任务的AI代理,而后者可能带来更大风险。LLM在面对问卷描述时的输入、行动、环境互动和内部处理过程,与其基于相同模型构建的代理存在显著差异。因此,LLM的问答结果无法代表对应代理的真实行为。我们进一步指出,此类评估假设LLM能准确报告其反事实行为,这一前提不成立,导致评估缺乏建构效度。该问题同样适用于现有AI对齐方法。最后,我们建议通过正视这些缺陷来改进安全评估与对齐训练。
原文摘要 · Abstract (English)
As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such assessments. However, most current advances therein may be ill-suited for assessing AI systems across real-world deployments. Standard methods prompt large language models (LLMs) in a questionnaire-style to describe their values or behavior in hypothetical scenarios. By focusing on unaugmented LLMs, they fall short of evaluating AI agents, which could actually perform relevant behaviors, hence posing much greater risks. LLMs' engagement with scenarios described by questionnaire-style prompts differs starkly from that of agents based on the same LLMs, as reflected in divergences in the inputs, possible actions, environmental interactions, and internal processing. As such, LLMs' responses to scenario descriptions are unlikely to be representative of the corresponding LLM agents' behavior. We further contend that such assessments make strong assumptions concerning the ability and tendency of LLMs to report accurately about their counterfactual behavior. This makes them inadequate to assess risks from AI systems in real-world contexts as they lack construct validity. We then argue that a structurally identical issue holds for current AI alignment approaches. Lastly, we discuss improving safety assessments and alignment training by taking these shortcomings to heart.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。