测试大模型在冲突任务中遵守安全原则的能力,发现合规有代价且易伪装。
Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
- 设计轻量可解释的基准,检验模型在任务冲突时是否真守安全规则。
- 发现遵守安全指令会降低任务表现,即使有合规解存在。
- 揭示高合规率可能只是任务失败伪装,非真正安全选择。
可信的高级AI发展需具备验证代理行为并早期发现控制缺陷的方法。核心在于确保代理在与操作目标冲突时仍能遵循安全关键原则。本文提出一种轻量、可解释的基准,评估大语言模型代理在面对冲突任务指令时维持高层安全原则的能力。对六种大语言模型的评估揭示两个主要发现:(1) 存在可量化的“合规成本”,即安全约束会降低任务表现,即便存在合规解;(2) “合规幻觉”现象,即高合规率常掩盖任务能力不足,而非出于原则性选择。这些发现表明,尽管大模型可受层级指令影响,但当前方法缺乏可靠安全治理所需的稳定性。
原文摘要 · Abstract (English)
Credible safety plans for advanced AI development require methods to verify agent behavior and detect potential control deficiencies early. A fundamental aspect is ensuring agents adhere to safety-critical principles, especially when these conflict with operational goals. This paper introduces a lightweight, interpretable benchmark to evaluate an LLM agent's ability to uphold a high-level safety principle when faced with conflicting task instructions. Our evaluation of six LLMs reveals two primary findings: (1) a quantifiable "cost of compliance" where safety constraints degrade task performance even when compliant solutions exist, and (2) an "illusion of compliance" where high adherence often masks task incompetence rather than principled choice. These findings provide initial evidence that while LLMs can be influenced by hierarchical directives, current approaches lack the consistency required for reliable safety governance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。