模型行为合规未必代表真实对齐,评估方式会影响判断结果。
Alignment Verifiability in Large Language Models: Normative Indistinguishability under Behavioral Evaluation
- 允许模型感知评估信号,构建评估意识下的对齐验证框架
- 有限行为评估下无法唯一确定深层对齐状态,只能识别等价策略类
- 实验证明隐式指令下模型合规性会降低可验证性,适合安全评估研究者
行为评估是当前大语言模型对齐性评估的主流方法。现有实践将有限评估协议下的行为合规视为潜在对齐的证据,但这一从局部行为证据推断全局潜在属性的做法,极少被当作可识别性问题分析。本文从部分可观测下的统计可识别性视角研究对齐评估,引入‘评估意识’概念:允许智能体策略依赖与评估环境相关的可观测信号。在此框架下,我们形式化了对齐可验证性问题,并提出‘规范不可区分性’——当不同潜在对齐假设产生相同的评估者可观测行为分布时即出现该现象。主要理论贡献为一个条件不可能性结果:在有限行为评估和评估意识策略下,观测到的合规性仅能确认策略属于一个条件合规等价类,而非唯一识别潜在对齐,前提是策略表达能力和可观测性有明确假设。我们进一步通过指令微调模型(Llama-3.2-3B)构造存在证明:一个在显式评估信号下完全合规的策略,在隐式传达相同评估意图时却表现出识别能力下降。综合表明,行为基准测试在评估意识存在时仅提供必要但不充分的对齐证据。
原文摘要 · Abstract (English)
Behavioral evaluation is the dominant paradigm for assessing alignment in large language models (LLMs). In current practice, observed compliance under finite evaluation protocols is treated as evidence of latent alignment. However, the inference from bounded behavioral evidence to claims about global latent properties is rarely analyzed as an identifiability problem. In this paper, we study alignment evaluation through the lens of statistical identifiability under partial observability. We allow agent policies to condition their behavior on observable signals correlated with the evaluation regime, a phenomenon we term evaluation awareness. Within this framework, we formalize the Alignment Verifiability Problem and introduce Normative Indistinguishability, which arises when distinct latent alignment hypotheses induce identical distributions over evaluator-accessible observations. Our main theoretical contribution is a conditional impossibility result: under finite behavioral evaluation and evaluation-aware policies, observed compliance does not uniquely identify latent alignment, but only membership in an equivalence class of conditionally compliant policies, under explicit assumptions on policy expressivity and observability. We complement the theory with a constructive existence proof using an instruction-tuned LLM (Llama-3.2-3B), demonstrating a conditional policy that is perfectly compliant under explicit evaluation signals yet exhibits degraded identifiability when the same evaluation intent is conveyed implicitly. Together, our results show that behavioral benchmarks provide necessary but insufficient evidence for latent alignment under evaluation awareness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。