测试大模型在多轮对话中遵守定制行为规范的能力,发现其合规性随对话轮次显著下降。
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
- 构建动态评估框架,模拟多轮对抗性场景测试模型对定制行为政策的遵循
- 开放与闭源模型在单轮中失败率低于4%,多轮对抗下最高达84%失败
- 适用于需要适配企业政策、行业规范等多元价值观的LLM应用研究
大型语言模型(LLMs)通常对通用安全和使用原则进行对齐,以适应广泛公众需求。然而,现实中的应用常发生在由企业政策、监管要求、使用场景、品牌指南和伦理承诺塑造的组织生态中。这凸显了对具备多元对齐目标的LLM进行严格、全面评估的需求,即强调模型对多样用户价值和需求的适应能力。本文提出PLURALISTIC BEHAVIOR SUITE(PBSUITE),一个动态评估套件,用于系统评估LLMs在多轮交互对话中遵守多元对齐规范的能力。PBSUITE包含:(1) 基于30个行业的300个真实世界行为政策的数据集;(2) 用于在对抗条件下压力测试模型对自定义行为规范遵从性的动态评估框架。实验发现,主流开源与闭源模型在单轮设置中表现出稳健的行为政策遵循(失败率低于4%),但在多轮对抗性交互中,合规性显著下降(最高失败率达84%)。结果表明,现有模型对齐与安全调控方法在真实场景中难以持续贯彻多元行为政策。本工作贡献了数据集与分析框架,支持未来面向鲁棒且情境感知的多元对齐技术研究。
原文摘要 · Abstract (English)
Large language models (LLMs) are typically aligned to a universal set of safety and usage principles intended for broad public acceptability. Yet, real-world applications of LLMs often take place within organizational ecosystems shaped by distinctive corporate policies, regulatory requirements, use cases, brand guidelines, and ethical commitments. This reality highlights the need for rigorous and comprehensive evaluation of LLMs with pluralistic alignment goals, an alignment paradigm that emphasizes adaptability to diverse user values and needs. In this work, we present PLURALISTIC BEHAVIOR SUITE (PBSUITE), a dynamic evaluation suite designed to systematically assess LLMs' capacity to adhere to pluralistic alignment specifications in multi-turn, interactive conversations. PBSUITE consists of (1) a diverse dataset of 300 realistic LLM behavioral policies, grounded in 30 industries; and (2) a dynamic evaluation framework for stress-testing model compliance with custom behavioral specifications under adversarial conditions. Using PBSUITE, We find that leading open- and closed-source LLMs maintain robust adherence to behavioral policies in single-turn settings (less than 4% failure rates), but their compliance weakens substantially in multi-turn adversarial interactions (up to 84% failure rates). These findings highlight that existing model alignment and safety moderation methods fall short in coherently enforcing pluralistic behavioral policies in real-world LLM interactions. Our work contributes both the dataset and analytical framework to support future research toward robust and context-aware pluralistic alignment techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。