通过对比分析检测大模型中的私有对齐行为
Auditing Proprietary Alignment in Large Language Models: A Comparative Framework Without a Ground-Truth Standard
- 基于对比行为分析,量化目标模型与基线模型的响应差异
- 无需真值标准,即可识别模型在敏感话题上的策略性偏差
- 适合研究者和监管方评估大模型背后的隐性政策
大型语言模型(LLMs)正通过不透明的开发与部署流程发布和应用,使模型提供方能够隐秘地注入特定于自身的政策而不公开声明。由此导致多个模型在争议性话题上生成反映机构利益或引发审查、误导性信息的内容。然而,系统性识别此类对齐行为仍面临根本挑战,尤其在于‘私有’在不同语境下的模糊性。本文提出一种统计框架,通过比较行为分析在黑盒条件下检测大模型中的私有对齐。该方法在共享语义空间中量化目标模型与一组基准模型响应之间的系统性偏离。通过评估相对行为差异而非绝对正确性,该框架实现了在黑盒访问下的合理审计。应用于若干此前未被量化的典型案例,为外部评估大模型中提供方特异性对齐行为提供了系统且可扩展的基础。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly released and deployed through opaque development and deployment pipelines, enabling model providers to inject intentional, provider-specific policies without officially announcing them. As a result, various models have been reported to generate responses reflecting proprietary rules and organizational interests, leading to censorship or misinformation on controversial topics. However, systematic identification of such alignment remains a fundamental challenge, complicated by the ambiguity of what ``proprietary'' entails in different contexts. In this paper, we propose a statistical framework for detecting proprietary alignment in black-box language models via comparative behavioral analysis. Our approach quantifies systematic deviations between the responses of a target model and those of a reference set of baseline models in a shared semantic space. By evaluating relative behavioral divergence rather than absolute correctness, our framework enables principled auditing under black-box access. Applied to several widely discussed but previously unquantified cases, it provides a systematic and scalable basis for external assessment of provider-specific alignment behavior in large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。