arXiv:2601.17087cs.HCcs.AI2026-01ACL被引 30

用大模型模拟用户评估智能体不可靠,尤其对非标准英语使用者偏差严重。

Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations

  • 用不同LLM模拟用户,同一智能体成功率相差达9个百分点。
  • 模拟用户低估难任务表现,高估中等难度任务,存在系统性偏差。
  • 对非裔美式英语和印度英语使用者效果更差,适合关注公平性的研究者参考。

代理评估基准越来越多地依赖大语言模型(LLM)模拟用户以可扩展地评估智能体性能,但该方法的鲁棒性、有效性与公平性尚未被检验。通过在美国、印度、肯尼亚和尼日利亚开展用户研究,我们考察了在τ-Bench零售任务中,LLM模拟用户是否能作为真实人类用户的可靠代理。结果发现,用户模拟缺乏鲁棒性,不同用户LLM下的智能体成功率达到9个百分点的差异。此外,使用模拟用户进行评估存在系统性校准偏差:在困难任务上低估智能体表现,在中等难度任务上则高估。非裔美式英语(AAVE)使用者的成功率始终更低,且误差随年龄增长显著加剧。模拟用户对不同人群的代理效果不一,尤其对AAVE和印度英语使用者最差。同时,模拟用户引入了对话层面的伪影,并暴露出不同于真人用户的问题模式。这些发现表明,当前评估实践可能扭曲智能体在多元用户群体中的真实能力,掩盖实际部署中的挑战。

原文摘要 · Abstract (English)

Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. Through a user study with participants across the United States, India, Kenya, and Nigeria, we investigate whether LLM-simulated users serve as reliable proxies for real human users in evaluating agents on τ-Bench retail tasks. We find that user simulation lacks robustness, with agent success rates varying up to 9 percentage points across different user LLMs. Furthermore, evaluations using simulated users exhibit systematic miscalibration, underestimating agent performance on challenging tasks and overestimating it on moderately difficult ones. African American Vernacular English (AAVE) speakers experience consistently worse success rates and calibration errors than Standard American English (SAE) speakers, with disparities compounding significantly with age. We also find simulated users to be a differentially effective proxy for different populations, performing worst for AAVE and Indian English speakers. Additionally, simulated users introduce conversational artifacts and surface different failure patterns than human users. These findings demonstrate that current evaluation practices risk misrepresenting agent capabilities across diverse user populations and may obscure real-world deployment challenges.

智能体评估公平性语言偏差模拟用户

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。