arXiv:2507.12284cs.SEcs.AI2025-07被引 5

首个面向非英语代码生成的多任务评估框架,覆盖8种语言11项任务。

MERA Code: A Unified Framework for Evaluating Code Generation Across Tasks

  • 构建涵盖8种语言11项任务的评估体系,聚焦真实编码能力。
  • 发现主流模型在非英语场景下执行实际编程任务能力显著不足。
  • 开源代码库+评分系统+排行榜,支持跨平台评测与持续对比。

大语言模型在软件工程中的自动化能力不断提升,但现有评估多集中于自然语言任务,忽视代码质量。多数基准侧重高层次推理而非可执行代码和真实场景表现,难以揭示模型在生产环境中的真实能力与风险。为此,我们提出MERA Code——MERA基准家族的新成员,专门针对俄语环境下最新代码生成大模型的评估。该基准包含11个评估任务,覆盖8种编程语言。其方法论基于一个实用编码技能分类体系,确保任务设计贴近真实开发需求。基准提供开源代码库,支持多种编程环境的评分系统,并配备排行榜与提交平台。我们评估了开源模型与前沿API模型在非英语编程任务中的表现,揭示其在实际应用中的局限性。MERA Code已公开发布,旨在引导未来研究、预判模型新特性并统一评估标准。

原文摘要 · Abstract (English)

Advancements in LLMs have enhanced task automation in software engineering; however, current evaluations primarily focus on natural language tasks, overlooking code quality. Most benchmarks prioritize high-level reasoning over executable code and real-world performance, leaving gaps in understanding true capabilities and risks associated with these models in production. To address this issue, we propose MERA Code, a new addition to the MERA benchmark family, specifically focused on evaluating code for the latest code generation LLMs in Russian. This benchmark includes 11 evaluation tasks that span 8 programming languages. Our proposed evaluation methodology features a taxonomy that outlines the practical coding skills necessary for models to complete these tasks. The benchmark comprises an open-source codebase for users to conduct MERA assessments, a scoring system compatible with various programming environments, and a platform featuring a leaderboard and submission system. We evaluate open LLMs and frontier API models, analyzing their limitations in terms of practical coding tasks in non-English languages. We are publicly releasing MERA to guide future research, anticipate groundbreaking features in model development, and standardize evaluation procedures.

代码生成多语言评估基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。