长上下文下大模型代理会突然失效,安全机制变得不可靠。
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
- 测试长上下文下智能体行为,发现性能与拒绝有害请求的能力随上下文长度剧烈波动。
- 10万token时性能下降超50%,部分模型拒绝率在20万token时翻了数倍。
- 提醒研究者:当前评估标准对长任务智能体不适用,需重新设计安全测试方法。
解决复杂或长周期问题通常需要大语言模型(LLMs)使用外部工具并在显著更长的上下文窗口中运行。新模型支持更长的上下文和工具调用能力。以往研究主要关注长上下文提示下的模型评估,对智能体设置在能力和安全方面的探索仍不足。本文填补这一空白。我们发现,LLM智能体对上下文长度、类型和位置极为敏感,任务表现和拒绝执行有害请求的能力出现意外且不一致的波动。拥有100万至200万标记上下文窗口的模型在达到10万标记时即出现严重性能下降,良性与有害任务的性能降幅均超过50%。拒绝率变化无常:GPT-4.1-nano的拒绝率从约5%升至约40%,而Grok 4 Fast则从约80%降至约10%(在20万标记处)。结果表明,长上下文运行的智能体存在潜在安全风险,并对现有评估指标和范式提出挑战。特别地,智能体在能力和安全表现上与以往针对相似条件的模型评估结果存在显著差异。
原文摘要 · Abstract (English)
Solving complex or long-horizon problems often requires large language models (LLMs) to use external tools and operate over a significantly longer context window. New LLMs enable longer context windows and support tool calling capabilities. Prior works have focused mainly on evaluation of LLMs on long-context prompts, leaving agentic setup relatively unexplored, both from capability and safety perspectives. Our work addresses this gap. We find that LLM agents could be sensitive to length, type, and placement of the context, exhibiting unexpected and inconsistent shifts in task performance and in refusals to execute harmful requests. Models with 1M-2M token context windows show severe degradation already at 100K tokens, with performance drops exceeding 50\% for both benign and harmful tasks. Refusal rates shift unpredictably: GPT-4.1-nano increases from $\sim$5\% to $\sim$40\% while Grok 4 Fast decreases from $\sim$80\% to $\sim$10\% at 200K tokens. Our work shows potential safety issues with agents operating on longer context and opens additional questions on the current metrics and paradigm for evaluating LLM agent safety on long multi-step tasks. In particular, our results on LLM agents reveal a notable divergence in both capability and safety performance compared to prior evaluations of LLMs on similar criteria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。