arXiv:2606.11070cs.CLcs.AI2026-06

构建跨领域多场景代理评估基准,提升真实复杂任务测试能力

T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains

论文配图:T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains
图 1 · 摘自论文原文
  • 设计25个跨领域交错场景,模拟真实多步交互与协同推理
  • 覆盖12个开源及私有模型,验证工具调用与对话质量的综合表现
  • 融合自动评估与人工打分,适合研究智能体系统与人机交互的学者

大语言模型在推理与工具调用方面的发展推动了智能体系统的进步。然而,现有基准在任务复杂度、现实性和领域多样性上仍显不足,难以捕捉跨多个领域的交互,限制了对需持续推理与协调的多步骤场景的评估。为此,我们提出T1-Bench,一个高保真、全面的基准,用于评估客户导向的多领域环境中的智能体系统。该基准包含25个不同难度的领域,通过交织的场景设计,要求在多轮用户-助手交互中进行结构化推理,显著提升了任务的组合复杂度与评估严谨性。我们使用12个开源及专有模型对T1-Bench进行了评估,提供可复现的标准框架,以衡量代理行为、工具使用与对话质量。此外,结合自动评估与人工判断,增强对定性性能的评估。T1-Bench通过提升任务复杂度、交互深度和领域覆盖,显著超越了以往基准。为促进未来研究,我们将公开数据与评估代码。

原文摘要 · Abstract (English)

Recent advances in reasoning and tool-calling capabilities of large language models (LLMs) have enabled increasingly capable agentic systems. However, existing benchmarks remain limited in task complexity, realism, and domain diversity, and often fail to capture interactions that span multiple domains, limiting their ability to evaluate agents in realistic multi-step settings that require sustained reasoning and coordination. To address these limitations, we introduce T1-Bench, a high-fidelity, comprehensive benchmark for evaluating agentic systems in realistic customer-facing, multi-domain environments, featuring interleaved scenarios that require structured reasoning across multi-turn user-assistant interactions and substantially increasing both compositional complexity and evaluative rigor across 25 domains of varying difficulty. We evaluate T1-Bench using 12 proprietary and open-weight models, providing a reproducible and standardized framework for assessing agent behavior, tool utilization, and conversational quality in complex, multi-step environments. We further complement automatic evaluation with human judgments to strengthen the assessment of qualitative performance. Overall, T1-Bench substantially advances prior benchmarks by increasing task complexity, interaction depth, and domain coverage in simulated multi-domain environments. To facilitate future research on agentic systems, we will publicly release data and evaluation code as open source.

智能体评估多领域对话系统基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。