arXiv:2505.16667cs.AI2025-05ACL被引 17

首个面向人-大模型编程协作的综合性评测基准,系统评估协同解题能力。

ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming

  • 构建人类反馈分类体系,细化编程全流程评估维度。
  • 推出可模拟真人反馈的ELABORATIONSET数据集,支持大规模测试。
  • 提供全面评测框架,揭示现有方法优缺点,指导未来研究。

尽管近期研究日益重视人与大模型在编程竞赛中的协作价值,并提出多种实证方法,但因现有研究分散且使用多样化的特定应用场景人类反馈,整体理解仍不清晰。为此,本文实现三个目标:第一,提出首个涵盖完整编程过程的人类反馈分类体系,支持细粒度评估;第二,构建ELABORATIONSET——一个专为人-大模型协作设计的编程数据集,经精心标注,支持大规模模拟人类反馈及低成本真实人类交互研究;第三,推出ELABORATION评测基准,实现对人-大模型编程协作能力的全面评估。通过该基准,我们识别出当前方法的优势与不足,为后续改进奠定基础。代码与数据集已开源至https://github.com/SCUNLP/ELABORATION。

原文摘要 · Abstract (English)

While recent research increasingly emphasizes the value of human-LLM collaboration in competitive programming and proposes numerous empirical methods, a comprehensive understanding remains elusive due to the fragmented nature of existing studies and their use of diverse, application-specific human feedback. Thus, our work serves a three-fold purpose: First, we present the first taxonomy of human feedback consolidating the entire programming process, which promotes fine-grained evaluation. Second, we introduce ELABORATIONSET, a novel programming dataset specifically designed for human-LLM collaboration, meticulously annotated to enable large-scale simulated human feedback and facilitate costeffective real human interaction studies. Third, we introduce ELABORATION, a novel benchmark to facilitate a thorough assessment of human-LLM competitive programming. With ELABORATION, we pinpoint strengthes and weaknesses of existing methods, thereby setting the foundation for future improvement. Our code and dataset are available at https://github.com/SCUNLP/ELABORATION

编程竞赛人机协作评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。