arXiv:2602.12285cs.CLcs.AI2026-02中稿 · AAAI被引 3

身份设定会严重削弱大模型代理的可靠性,最高导致26.2%性能下降。

From Biased Chatbots to Biased Agents: Examining Role Assignment Effects on LLM Agent Robustness

  • 通过角色设定测试大模型代理在不同任务中的表现
  • 发现身份设定导致性能最高下降26.2%,且影响跨任务与模型架构
  • 警示身份设定可能引入隐性偏见,适合关注安全部署的研究者阅读

大型语言模型(LLMs)正越来越多地作为具备真实世界影响的自主代理使用,而不仅仅是文本生成。尽管角色诱导的文本生成偏见已有研究,但其对代理任务表现的影响仍基本未被探索,尽管此类影响带来更直接的操作风险。本文首次系统性地揭示:基于人口统计学的角色设定可改变大模型代理的行为,并在多个领域内降低其性能。在涵盖战略推理、规划与技术操作的多类代理基准上评估广泛部署的模型,发现高达26.2%的性能下降,由与任务无关的角色线索引发。这种变化出现在各类任务和模型架构中,表明角色条件化与简单提示注入可扭曲代理决策的可靠性。研究揭示当前大模型代理系统的潜在漏洞:角色设定可能引入隐性偏见并增加行为波动性,对大模型代理的安全与鲁棒部署构成严峻挑战。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of actions with real-world impacts beyond text generation. While persona-induced biases in text generation are well documented, their effects on agent task performance remain largely unexplored, even though such effects pose more direct operational risks. In this work, we present the first systematic case study showing that demographic-based persona assignments can alter LLM agents' behavior and degrade performance across diverse domains. Evaluating widely deployed models on agentic benchmarks spanning strategic reasoning, planning, and technical operations, we uncover substantial performance variations - up to 26.2% degradation, driven by task-irrelevant persona cues. These shifts appear across task types and model architectures, indicating that persona conditioning and simple prompt injections can distort an agent's decision-making reliability. Our findings reveal an overlooked vulnerability in current LLM agentic systems: persona assignments can introduce implicit biases and increase behavioral volatility, raising concerns for the safe and robust deployment of LLM agents.

大模型代理角色偏见性能下降安全部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。