用可调控用户特质模拟测试AI代理鲁棒性,发现当前模型易因用户急躁等行为失效。
Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
- 通过激活空间方向控制用户特质(如急躁、混乱),无需微调即可生成多样化用户行为。
- 在四个真实场景中,前沿模型性能平均下降2%-30%,暴露其对用户行为变化的脆弱性。
- 开源工具支持社区开展逼真、多变的代理质量评估,适合开发者和评测人员使用。
尽管对话式AI代理发展迅速,其鲁棒性仍缺乏有效测试。用户行为的微小变化,如更急躁、不连贯或怀疑态度,会导致代理性能急剧下降,暴露出当前AI代理的脆弱性。现有基准无法捕捉这种脆弱性:代理在标准评估中表现良好,但在更真实多样的环境中却严重退化。为此,我们提出TraitBasis,一种轻量级、模型无关的方法,用于系统性压力测试AI代理。TraitBasis学习激活空间中对应可调控用户特质(如急躁或不连贯)的方向,可在推理时控制、缩放、组合使用,无需微调或额外数据。利用TraitBasis,我们将τ-Bench扩展为τ-Trait,通过受控特质向量改变用户行为。在主流模型上,τ-Trait平均导致2%-30%的性能下降,凸显当前代理对用户行为变化的鲁棒性不足。这些结果强调了鲁棒性测试的重要性,也展示了TraitBasis作为简单、高效、可组合工具的潜力。通过支持模拟驱动的压力测试与训练循环,TraitBasis为构建能在真实人类互动中保持可靠的代理铺平道路。我们已将τ-Trait在航空、零售、电信和远程医疗四个领域开源,供社区在真实、行为多样的情境下系统性验证代理性能:https://github.com/collinear-ai/tau-trait。
原文摘要 · Abstract (English)
Despite rapid progress in building conversational AI agents, robustness is still largely untested. Small shifts in user behavior, such as being more impatient, incoherent, or skeptical, can cause sharp drops in agent performance, revealing how brittle current AI agents are. Today's benchmarks fail to capture this fragility: agents may perform well under standard evaluations but degrade spectacularly in more realistic and varied settings. We address this robustness testing gap by introducing TraitBasis, a lightweight, model-agnostic method for systematically stress testing AI agents. TraitBasis learns directions in activation space corresponding to steerable user traits (e.g., impatience or incoherence), which can be controlled, scaled, composed, and applied at inference time without any fine-tuning or extra data. Using TraitBasis, we extend $τ$-Bench to $τ$-Trait, where user behaviors are altered via controlled trait vectors. We observe on average a 2%-30% performance degradation on $τ$-Trait across frontier models, highlighting the lack of robustness of current AI agents to variations in user behavior. Together, these results highlight both the critical role of robustness testing and the promise of TraitBasis as a simple, data-efficient, and compositional tool. By powering simulation-driven stress tests and training loops, TraitBasis opens the door to building AI agents that remain reliable in the unpredictable dynamics of real-world human interactions. We have open-sourced $τ$-Trai across four domains: airline, retail, telecom, and telehealth, so the community can systematically QA their agents under realistic, behaviorally diverse intents and trait scenarios: https://github.com/collinear-ai/tau-trait.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。