arXiv:2506.13901cs.CLcs.AI2025-06EMNLP被引 6

用隐空间几何分析模型对齐质量,发现拒绝率之外的隐藏风险

Alignment Quality Index (AQI) : Beyond Refusals: AQI as an Intrinsic Alignment Diagnostic via Latent Geometry, Cluster Divergence, and Layer wise Pooled Representations

  • 通过激活聚类分离度评估对齐,不依赖输出行为
  • 在多个模型上验证了与人工判断的高度相关性
  • 适合安全审计、对抗攻击检测及对齐研究者使用

随着大语言模型进入教育、医疗、治理等高风险领域,其行为必须可靠地符合人类价值观与安全约束。当前评估主要依赖拒绝率、G-Eval分数和毒性分类器等行为代理,存在明显盲区:对齐模型仍可能被越狱、受生成随机性影响或出现对齐伪装。为此,我们提出对齐质量指数(AQI),一种基于隐空间几何的提示无关度量方法,通过分析安全与不安全激活在潜在空间中的分离程度来评估对齐。结合戴维斯-鲍尔丁指数(DBS)、邓恩指数(DI)、谢-贝尼指数(XBI)和卡林斯基-哈拉巴兹指数(CHI)等多种指标,AQI能有效捕捉聚类质量,揭示隐藏的错位与越狱风险,即使输出看似合规。此外,它还能作为对齐伪装的早期预警信号,提供解码无关、行为无偏的安全审计工具。我们还构建了LITMUS数据集以支持在此挑战条件下的稳健评估。在多种基于DPO、GRPO和RLHF训练的模型上,实证测试显示AQI与外部评委评分高度相关,并能发现拒绝率遗漏的漏洞。代码已公开,推动该领域研究。

原文摘要 · Abstract (English)

Alignment is no longer a luxury, it is a necessity. As large language models (LLMs) enter high-stakes domains like education, healthcare, governance, and law, their behavior must reliably reflect human-aligned values and safety constraints. Yet current evaluations rely heavily on behavioral proxies such as refusal rates, G-Eval scores, and toxicity classifiers, all of which have critical blind spots. Aligned models are often vulnerable to jailbreaking, stochasticity of generation, and alignment faking. To address this issue, we introduce the Alignment Quality Index (AQI). This novel geometric and prompt-invariant metric empirically assesses LLM alignment by analyzing the separation of safe and unsafe activations in latent space. By combining measures such as the Davies-Bouldin Score (DBS), Dunn Index (DI), Xie-Beni Index (XBI), and Calinski-Harabasz Index (CHI) across various formulations, AQI captures clustering quality to detect hidden misalignments and jailbreak risks, even when outputs appear compliant. AQI also serves as an early warning signal for alignment faking, offering a robust, decoding invariant tool for behavior agnostic safety auditing. Additionally, we propose the LITMUS dataset to facilitate robust evaluation under these challenging conditions. Empirical tests on LITMUS across different models trained under DPO, GRPO, and RLHF conditions demonstrate AQI's correlation with external judges and ability to reveal vulnerabilities missed by refusal metrics. We make our implementation publicly available to foster future research in this area.

对齐评估隐空间分析安全审计模型测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。