arXiv:2504.20673cs.SEcs.AI2025-04被引 1

构建多任务代码评估基准,全面测试大模型编程能力

CoCo-Bench: A Comprehensive Code Benchmark For Multi-task Large Language Model Evaluation

  • 设计覆盖理解、生成、修改、评审四维度的综合评测框架
  • 支持多语言和多难度任务,经人工审核确保数据质量
  • 揭示模型在真实开发场景中的性能差异,适合研究与工程应用

大型语言模型(LLMs)在软件工程中扮演关键角色,擅长代码生成与维护等任务。然而,现有评测基准普遍聚焦单一任务,缺乏反映真实应用场景的综合性评估体系。为此,我们提出 CoCo-Bench(Comprehensive Code Benchmark),涵盖代码理解、生成、修改和评审四大核心维度,全面覆盖开发者实际需求,实现系统化、代表性评估。该基准支持多种编程语言与不同难度的任务,并通过严格的人工审核保障数据准确性和质量。实验结果表明,CoCo-Bench 与现有基准具有一致性,同时揭示了模型在不同任务上的显著性能差异,有效凸显其优势与短板。该基准为代码导向型 LLM 的研究与技术发展提供客观、全面的评估依据,推动领域标准化进程。

原文摘要 · Abstract (English)

Large language models (LLMs) play a crucial role in software engineering, excelling in tasks like code generation and maintenance. However, existing benchmarks are often narrow in scope, focusing on a specific task and lack a comprehensive evaluation framework that reflects real-world applications. To address these gaps, we introduce CoCo-Bench (Comprehensive Code Benchmark), designed to evaluate LLMs across four critical dimensions: code understanding, code generation, code modification, and code review. These dimensions capture essential developer needs, ensuring a more systematic and representative evaluation. CoCo-Bench includes multiple programming languages and varying task difficulties, with rigorous manual review to ensure data quality and accuracy. Empirical results show that CoCo-Bench aligns with existing benchmarks while uncovering significant variations in model performance, effectively highlighting strengths and weaknesses. By offering a holistic and objective evaluation, CoCo-Bench provides valuable insights to guide future research and technological advancements in code-oriented LLMs, establishing a reliable benchmark for the field.

代码生成大模型评测软件工程多任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。