arXiv:2505.13195cs.AI2025-05被引 5

测试大模型在对抗环境下的决策弱点,揭示其易被操纵的隐患。

Adversarial Testing in LLMs: Insights into Decision-Making Vulnerabilities

  • 设计交互式对抗框架,模拟真实动态环境下的决策压力。
  • 多模型对比发现,不同模型在策略适应上存在显著差异。
  • 适合关注AI安全与对齐的研究者,推动更稳健的智能体设计。

随着大语言模型(LLMs)越来越多地应用于实际决策系统,理解其行为脆弱性成为AI安全与对齐的关键挑战。现有评估指标主要关注推理准确率或事实正确性,却常忽视模型在对抗性干扰下的鲁棒性,以及在动态环境中采用自适应策略的能力。本文提出一种对抗性评估框架,系统性地在交互式对抗条件下测试LLMs的决策过程。借鉴认知心理学与博弈论方法,该框架在两类经典任务中检验模型表现:两臂赌博机任务与多轮信任任务,分别捕捉探索-利用权衡、社会合作与策略灵活性等核心特征。我们对GPT-3.5、GPT-4、Gemini-1.5和DeepSeek-V3等先进模型进行测试,揭示了模型特有的易操控性及策略适应僵化问题。研究发现模型间行为模式差异显著,强调了适应能力与公平性认知对可信AI部署的重要性。本工作不提供性能基准,而是提出一种诊断LLM决策缺陷的方法,为对齐与安全研究提供可操作洞见。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) become increasingly integrated into real-world decision-making systems, understanding their behavioural vulnerabilities remains a critical challenge for AI safety and alignment. While existing evaluation metrics focus primarily on reasoning accuracy or factual correctness, they often overlook whether LLMs are robust to adversarial manipulation or capable of using adaptive strategy in dynamic environments. This paper introduces an adversarial evaluation framework designed to systematically stress-test the decision-making processes of LLMs under interactive and adversarial conditions. Drawing on methodologies from cognitive psychology and game theory, our framework probes how models respond in two canonical tasks: the two-armed bandit task and the Multi-Round Trust Task. These tasks capture key aspects of exploration-exploitation trade-offs, social cooperation, and strategic flexibility. We apply this framework to several state-of-the-art LLMs, including GPT-3.5, GPT-4, Gemini-1.5, and DeepSeek-V3, revealing model-specific susceptibilities to manipulation and rigidity in strategy adaptation. Our findings highlight distinct behavioral patterns across models and emphasize the importance of adaptability and fairness recognition for trustworthy AI deployment. Rather than offering a performance benchmark, this work proposes a methodology for diagnosing decision-making weaknesses in LLM-based agents, providing actionable insights for alignment and safety research.

大模型安全对抗测试决策脆弱性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。