arXiv:2512.04111cs.SEcs.AI2025-12中稿 · ICML被引 10

评测人机协作编程中人类价值,发现合作胜过单打独斗。

CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding

  • 设计需人机协作才能解的题目模板,逼出真实协同能力。
  • 人机合作通过率31.11%,远超单独人类(18.89%)和模型(0.67%)。
  • 适合研究人机协同、编程智能或智能开发工具的学者与工程师。

基于大模型的编程代理正在重塑开发范式。然而,现有评估体系既非传统的人类测试,也非专为大模型设计的基准,无法捕捉人机协作的转变——即问题需人类推理引导方向、AI高效执行实现。我们提出 CentaurEval,一个统一且生态有效的基准,用于衡量编程中的人机协同价值。其核心是45种‘协作必要’的问题模板,单独人类或模型均无法解决,但通过有效协作可完成。该基准动态生成任务,为人提供标准化开发环境,为模型提供可复现的450个任务工具包。我们在4种人类干预级别下对45名参与者和5个大模型进行评估。结果表明,单独人类通过率为18.89%,模型仅为0.67%,而人机协作提升至31.11%。分析揭示出一种新型共推理关系,挑战了传统的‘人类主导-工具辅助’层级,表明战略突破可能源自人类或AI任一方。

原文摘要 · Abstract (English)

LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to capture this shift, excluding problems that require both human reasoning to guide solutions and AI efficiency for implementation. We introduce CentaurEval, a unified, ecologically valid benchmark for measuring human-in-the-loop value in coding. CentaurEval's core innovation is its "Collaboration-Necessary" problem templates, which are intractable for standalone LLMs or humans, but solvable through effective collaboration. CentaurEval dynamically instantiates tasks from 45 templates, providing a standardized IDE for humans and a reproducible 450-task toolkit for LLMs. We benchmark 45 participants against 5 LLMs under 4 levels of human intervention. Results show that while LLMs or humans alone achieve poor pass rates (0.67% and 18.89%), human-AI collaboration significantly improves to 31.11%. Our analysis reveals an emerging co-reasoning partnership, challenging the traditional human-tool hierarchy by showing that strategic breakthroughs can originate from either humans or AI.

人机协作编程智能评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。