arXiv:2503.22115cs.CLcs.AI2025-03被引 2

用对话和故事测试大模型伦理对齐,更难伪装。

Beyond Single-Sentence Prompts: Upgrading Value Alignment Benchmarks with Dialogues and Stories

  • 用多轮对话和叙事场景替代单句提问,提升评估隐蔽性。
  • 新方法能暴露传统测试中未发现的潜在偏见。
  • 适合关注AI伦理与安全评估的研究者使用。

评估大语言模型(LLMs)的价值对齐通常依赖单句对抗性提示,直接询问具有伦理敏感性或争议性的问题。然而,随着AI安全技术的快速发展,模型已越来越擅长规避这些简单测试,导致其无法有效揭示潜在偏见和伦理立场。为此,我们提出一种升级版价值对齐基准,突破单句提示的局限,引入多轮对话与基于叙事的情境。该方法增强评估的隐蔽性与对抗性,使其更能抵御现代LLM中表面的安全防护。我们构建了一个包含对话陷阱与伦理模糊叙事的数据集,系统评估模型在更复杂、上下文丰富的场景下的响应。实验表明,该方法能有效暴露传统单次评估中未检测到的深层偏见。研究强调了在动态、情境化环境中进行价值对齐评估的重要性,为更先进、更真实的AI伦理与安全评估开辟了路径。

原文摘要 · Abstract (English)

Evaluating the value alignment of large language models (LLMs) has traditionally relied on single-sentence adversarial prompts, which directly probe models with ethically sensitive or controversial questions. However, with the rapid advancements in AI safety techniques, models have become increasingly adept at circumventing these straightforward tests, limiting their effectiveness in revealing underlying biases and ethical stances. To address this limitation, we propose an upgraded value alignment benchmark that moves beyond single-sentence prompts by incorporating multi-turn dialogues and narrative-based scenarios. This approach enhances the stealth and adversarial nature of the evaluation, making it more robust against superficial safeguards implemented in modern LLMs. We design and implement a dataset that includes conversational traps and ethically ambiguous storytelling, systematically assessing LLMs' responses in more nuanced and context-rich settings. Experimental results demonstrate that this enhanced methodology can effectively expose latent biases that remain undetected in traditional single-shot evaluations. Our findings highlight the necessity of contextual and dynamic testing for value alignment in LLMs, paving the way for more sophisticated and realistic assessments of AI ethics and safety.

伦理对齐评估基准对话测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。