通过对抗性多轮测试,发现大模型在持续攻击下的道德脆弱性。
Adversarial Moral Stress Testing of Large Language Models
- 设计多轮对抗压力测试框架,模拟真实恶意交互场景。
- 揭示模型在持续攻击下行为退化、尾部风险上升等隐藏缺陷。
- 适用于评估部署中AI系统的伦理鲁棒性,尤其适合高风险应用。
评估部署在软件系统中的大语言模型(LLMs)的伦理鲁棒性仍具挑战性,尤其是在持续对抗性用户交互下。现有安全基准通常依赖单轮评估和聚合指标(如毒性得分、拒绝率),难以捕捉真实多轮交互中可能出现的行为不稳定性。因此,罕见但高影响的伦理失败和渐进式退化现象可能在部署前未被发现。本文提出对抗性道德压力测试(AMST),一种基于压力的评估框架,用于在对抗性多轮交互下评估伦理鲁棒性。AMST对提示施加结构化压力变换,并通过分布感知的鲁棒性指标(包括方差、尾部风险、时序行为漂移)评估模型表现。我们在多个先进LLM上进行测试,包括LLaMA-3-8B、GPT-4o和DeepSeek-v3,使用受控压力条件下生成的大规模对抗场景。结果表明,不同模型的鲁棒性特征差异显著,暴露了传统单轮评估无法察觉的退化模式。特别地,鲁棒性更依赖于分布稳定性与尾部行为,而非平均性能。此外,AMST提供可扩展、模型无关的压力测试方法,支持对运行在对抗环境中的LLM系统进行鲁棒性感知评估与监控。
原文摘要 · Abstract (English)
Evaluating the ethical robustness of large language models (LLMs) deployed in software systems remains challenging, particularly under sustained adversarial user interaction. Existing safety benchmarks typically rely on single-round evaluations and aggregate metrics, such as toxicity scores and refusal rates, which offer limited visibility into behavioral instability that may arise during realistic multi-turn interactions. As a result, rare but high-impact ethical failures and progressive degradation effects may remain undetected prior to deployment. This paper introduces Adversarial Moral Stress Testing (AMST), a stress-based evaluation framework for assessing ethical robustness under adversarial multi-round interactions. AMST applies structured stress transformations to prompts and evaluates model behavior through distribution-aware robustness metrics that capture variance, tail risk, and temporal behavioral drift across interaction rounds. We evaluate AMST on several state-of-the-art LLMs, including LLaMA-3-8B, GPT-4o, and DeepSeek-v3, using a large set of adversarial scenarios generated under controlled stress conditions. The results demonstrate substantial differences in robustness profiles across models and expose degradation patterns that are not observable under conventional single-round evaluation protocols. In particular, robustness has been shown to depend on distributional stability and tail behavior rather than on average performance alone. Additionally, AMST provides a scalable and model-agnostic stress-testing methodology that enables robustness-aware evaluation and monitoring of LLM-enabled software systems operating in adversarial environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。