arXiv:2608.18554cs.CYcs.AI2026-08

测试大模型在辅助人类工作时的真实表现,发现能自动完成任务的模型未必能更好协助他人。

CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

论文配图:CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
图 1 · 摘自论文原文
  • 设计双模式评测框架:自动化与辅助增强并行测试
  • 7项真实任务中,仅1个模型平均优于无指导状态
  • 自动化表现强的模型在辅助场景中反而常落后

现有大模型评测多聚焦于自动完成任务的能力。但在实际应用中,模型更多用于辅助其他(人类或模型)代理。因此,选择模型不仅要看输出质量,更要看其能否提升弱代理的工作表现。本文提出统一评测框架,评估模型在自动化与辅助增强另一代理表现上的能力。在七项经济相关的现实任务中,助手模型为标准化低能力工作者生成辅助文本,由该工作者生成最终成果;在自动化模式下,助手直接生成输出。通过十次重复的盲评对比较,由特定任务评分标准的LLM评委打分。两种模式下的排名相关性较弱,自动化领先者在五项任务中反而在辅助模式中表现不佳。辅助效果并非总是正向:未受助的工作者在三项任务中高于所有辅助条件,且仅一个模型的指导在平均上优于无指导。结果表明,自动化能力不能完全代表辅助质量,亟需面向人机协同与多智能体系统中角色分工的评测基准。

原文摘要 · Abstract (English)

Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model's guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.

大模型评测人机协作辅助增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。