arXiv:2602.11348cs.AI2026-02被引 15

测试大模型智能体在噪声环境下的鲁棒性,发现真实场景下性能显著下降。

AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition

  • 构建噪声注入框架,模拟用户和工具两类真实噪声
  • 多模型测试显示噪声下性能普遍下降,暴露脆弱性
  • 适合关注智能体落地可靠性的研究者与工程师

大语言模型驱动的智能体在基准测试中表现优异,但在实际部署中常表现不佳,尤其在复杂不完美的环境中。这主要源于现有训练与评估范式依赖理想化假设,忽视了真实交互中的随机性和噪声。为此,我们提出AgentNoiseBench,一个系统评估智能体在噪声环境下鲁棒性的框架。通过深入分析真实场景中的偏差与不确定性,我们将环境噪声分为用户噪声和工具噪声两类。基于此,开发自动化管道,在保持任务可解性的前提下向现有代理基准注入可控噪声。利用该管道,对多种架构与参数规模的模型进行了广泛评估。结果表明,不同噪声条件下性能存在持续波动,揭示当前代理模型对真实环境扰动的高度敏感性。

原文摘要 · Abstract (English)

Recent advances in large language models have enabled LLM-based agents to achieve strong performance on a variety of benchmarks. However, their performance in real-world deployments often that observed on benchmark settings, especially in complex and imperfect environments. This discrepancy largely arises because prevailing training and evaluation paradigms are typically built on idealized assumptions, overlooking the inherent stochasticity and noise present in real-world interactions. To bridge this gap, we introduce AgentNoiseBench, a framework for systematically evaluating the robustness of agentic models under noisy environments. We first conduct an in-depth analysis of biases and uncertainties in real-world scenarios and categorize environmental noise into two primary types: user-noise and tool-noise. Building on this analysis, we develop an automated pipeline that injects controllable noise into existing agent-centric benchmarks while preserving task solvability. Leveraging this pipeline, we perform extensive evaluations across a wide range of models with diverse architectures and parameter scales. Our results reveal consistent performance variations under different noise conditions, highlighting the sensitivity of current agentic models to realistic environmental perturbations.

智能体鲁棒性噪声评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。