arXiv:2511.12306cs.AIcs.CY2025-11被引 4

用真实自由职业工作构建动态评估框架,测AI在真实职场中的表现。

UpBench: A Dynamically Evolving Real-World Labor-Market Agentic Benchmark Framework Built for Human-Centric AI

  • 基于全球自由职业平台真实任务构建动态评测体系。
  • 专家按详细标准评分,可分析AI的执行细节与指令遵循能力。
  • 适合研究人机协作、AI职场应用的学者与开发者。

随着大语言模型代理越来越多地承担数字工作,亟需可靠的评估框架来衡量其在真实世界中的能力、适应性及人机协作潜力。现有基准多为静态、合成或领域受限,难以反映代理在动态且具有经济意义环境下的表现。我们提出UpBench,一个基于全球自由职业平台Upwork上真实客户交易的任务基准。每个任务对应一次经验证的交易,确保评估建立在真实的职场活动与财务结果之上。UpBench采用基于评分标准的评估框架,由专业自由职业者将每项工作拆解为可验证的接受标准,并对AI提交内容逐项打分反馈,实现对模型优势、短板和指令遵循精度的细粒度分析,超越简单的通过/失败判断。人类专家贯穿数据流程(任务筛选、标准制定到评估),保障与真实职业标准一致,支持人机协作研究。通过定期更新任务以反映在线工作的持续演变,UpBench为智能体系统在真实劳动力市场情境下的评估提供了可扩展、以人为本的基础,推动构建以合作而非替代为核心的新型人机协同模式。

原文摘要 · Abstract (English)

As large language model (LLM) agents increasingly undertake digital work, reliable frameworks are needed to evaluate their real-world competence, adaptability, and capacity for human collaboration. Existing benchmarks remain largely static, synthetic, or domain-limited, providing limited insight into how agents perform in dynamic, economically meaningful environments. We introduce UpBench, a dynamically evolving benchmark grounded in real jobs drawn from the global Upwork labor marketplace. Each task corresponds to a verified client transaction, anchoring evaluation in genuine work activity and financial outcomes. UpBench employs a rubric-based evaluation framework, in which expert freelancers decompose each job into detailed, verifiable acceptance criteria and assess AI submissions with per-criterion feedback. This structure enables fine-grained analysis of model strengths, weaknesses, and instruction-following fidelity beyond binary pass/fail metrics. Human expertise is integrated throughout the data pipeline (from job curation and rubric construction to evaluation) ensuring fidelity to real professional standards and supporting research on human-AI collaboration. By regularly refreshing tasks to reflect the evolving nature of online work, UpBench provides a scalable, human-centered foundation for evaluating agentic systems in authentic labor-market contexts, offering a path toward a collaborative framework, where AI amplifies human capability through partnership rather than replacement.

AI评估人机协作真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。