arXiv:2603.07980cs.LGcs.AI2026-03被引 10

评测语言代理在真实专业场景中的表现,400个高难度任务挑战人类专家水平。

\$OneMillion-Bench: How Far are Language Agents from Human Experts?

  • 构建400个跨领域的专家级任务,涵盖法律、金融等真实场景。
  • 要求检索权威资料、处理矛盾证据,正确性依赖推理过程而非仅答案。
  • 适合评估代理在专业领域的真实可用性,推动智能体落地应用。

随着语言模型从聊天助手演变为具备多步推理和工具使用能力的长时程智能体,现有基准仍局限于结构化或考试式任务,难以反映真实职业需求。为此,我们提出$OneMillion-Bench,一个包含400个专家精选任务的基准,覆盖法律、金融、工业、医疗和自然科学领域,旨在评估智能体在具有经济影响的实际场景中的表现。与以往工作不同,该基准要求检索权威来源、解决冲突证据、应用领域特定规则并做出约束性决策,正确性既取决于最终答案,也取决于推理过程。我们采用基于评分标准的评估协议,衡量事实准确性、逻辑连贯性、实践可行性和专业合规性,聚焦专家级问题以确保智能体间可有效区分。$OneMillion-Bench为评估智能体在高专业度场景中的可靠性、深度和实用性提供了统一测试平台。

原文摘要 · Abstract (English)

As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce \$OneMillion-Bench \$OneMillion-Bench, a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focused on expert-level problems to ensure meaningful differentiation across agents. Together, \$OneMillion-Bench provides a unified testbed for assessing agentic reliability, professional depth, and practical readiness in domain-intensive scenarios.

语言智能体评测基准专业场景多步推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。