测试大模型行为规范,发现多处原则冲突与模糊地带。
Stress-Testing Model Specs Reveals Character Differences among Language Models
- 构建场景迫使模型在对立价值观间做选择,检测规范缺陷。
- 12个前沿模型在7万余个场景中表现严重分歧,暴露规范问题。
- 适合关注AI伦理对齐、安全评估的研究者与开发者参考。
大型语言模型(LLMs)越来越多地基于人工智能准则和模型规格进行训练,以确立行为规范与伦理原则。然而,这些规格面临内部原则冲突和复杂情境覆盖不足等关键挑战。本文提出一种系统性方法,用于压力测试模型行为规范,自动识别当前模型规格中存在的大量原则矛盾与解释模糊性。通过构建涵盖多元价值权衡的场景,强制模型在无法同时满足的合法原则之间做出抉择,我们对来自主要厂商(Anthropic、OpenAI、Google、xAI)的12个前沿模型进行了评估,采用价值分类得分衡量行为分歧。结果显示,在测试场景中发现了超过7万例显著的行为差异。实证表明,模型行为的高分歧性强烈预示着底层规范中的根本问题。定性分析揭示了当前模型规格中存在直接矛盾及若干原则的解释模糊性。此外,生成的数据集还暴露出所有模型中明显的对齐偏差与误拒情况。最后,我们梳理并比较了各模型的价值优先级模式与差异。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly trained from AI constitutions and model specifications that establish behavioral guidelines and ethical principles. However, these specifications face critical challenges, including internal conflicts between principles and insufficient coverage of nuanced scenarios. We present a systematic methodology for stress-testing model character specifications, automatically identifying numerous cases of principle contradictions and interpretive ambiguities in current model specs. We stress test current model specs by generating scenarios that force explicit tradeoffs between competing value-based principles. Using a comprehensive taxonomy we generate diverse value tradeoff scenarios where models must choose between pairs of legitimate principles that cannot be simultaneously satisfied. We evaluate responses from twelve frontier LLMs across major providers (Anthropic, OpenAI, Google, xAI) and measure behavioral disagreement through value classification scores. Among these scenarios, we identify over 70,000 cases exhibiting significant behavioral divergence. Empirically, we show this high divergence in model behavior strongly predicts underlying problems in model specifications. Through qualitative analysis, we provide numerous example issues in current model specs such as direct contradiction and interpretive ambiguities of several principles. Additionally, our generated dataset also reveals both clear misalignment cases and false-positive refusals across all of the frontier models we study. Lastly, we also provide value prioritization patterns and differences of these models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。