arXiv:2606.01815cs.CL2026-06

构建复杂任务依赖与真实用户模拟的评估基准,检验大模型在高难度场景下的推理能力。

CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation

论文配图:CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation
图 1 · 摘自论文原文
  • 基于约束图生成多实体依赖任务,含数千干扰项,仅少量解有效。
  • 顶尖模型在基准上通过率仅61%,引入真实用户模拟后下降最高达57%。
  • 揭示信息泄露最致命,模型更倾向隐匿错误而非承认失误。

在真实服务场景中评估大语言模型代理,需考虑复杂任务依赖、不完美用户行为及多种合理解的存在。我们提出CRAB-Bench(基于约束的真实代理评估基准)和RUSE(真实用户模拟引擎)以填补这一空白。CRAB-Bench通过多实体间的约束图生成任务,包含结构化干扰项,要求代理在数千个误导性候选中进行精细推理,其中只有极小部分为有效解。RUSE将协作性、模板化的模拟器替换为基于人类行为研究的真实用户,涵盖多样人格及四个行为维度。对四个前沿大模型的实验表明,最优模型在CRAB-Bench上的pass@1仅为61%,切换至RUSE后性能进一步下降高达57%,主要体现在任务求解能力而非对话质量。信息泄露是影响最严重的维度,且与RUSE交互的模型更少承认错误,倾向于通过隐式修正掩盖缺陷。

原文摘要 · Abstract (English)

Evaluating LLM agents in realistic service scenarios requires complex task dependencies, imperfect user behavior, and an evaluation that accommodates multiple valid solutions. We introduce CRAB-Bench (Constraint-based Realistic Agent Benchmark) and RUSE (Realistic User Simulation Engine) to address this gap. CRAB-Bench generates tasks via a constraint graph over multiple interdependent entities with structured distractors, requiring agents to reason carefully over thousands of misleading candidates where only a tiny fraction of solutions are valid. RUSE replaces cooperative, template-like simulators with realistic users grounded in human behavioral studies, instantiated across diverse personas and four behavioral dimensions. Experiments on four frontier LLM agents show that the best model achieves only 61% pass@1 on CRAB-Bench, and switching to RUSE causes further drops of up to 57%, concentrated in task-solving ability rather than conversational quality. Information Disclosure is the most damaging behavioral dimension, and agents interacting with RUSE are less likely to admit mistakes, instead masking errors through implicit corrections.

LLM评估任务依赖用户模拟推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。