构建首个覆盖6大印地语族语言的代码生成评测基准,推动多语言编程AI普惠化。
IndicEval-XL: Bridging Linguistic Diversity in Code Generation Across Indic Languages
- 构建跨6种印地语族语言与12种编程语言的代码生成评测框架
- 涵盖约14%全球人口使用的语言,提升开发工具的包容性
- 开源数据集和评测工具,助力多语言AI开发研究
大型语言模型在自然语言转代码方面展现出卓越能力,正重塑软件开发流程。然而现有评测基准仍以英语为主,难以覆盖全球开发者。为此,我们提出IndicEval-XL,一个面向6种主要印地语族语言的综合性代码生成评测基准,这些语言由约14%的全球人口使用。该基准将印地语族语言与12种编程语言对接,构建了稳健的评估体系。鉴于印度占全球人口八分之一,且印地语族语言在社会中具有关键作用,本工作对扩展代码生成系统的语言多样性具有重要意义。通过开发多语言支持资源,我们致力于让AI开发工具更包容、更易用。相关数据集与评测基准已开源,可访问 https://github.com/telekom/IndicEval-XL 获取。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation from natural language prompts, revolutionizing software development workflows. As we advance towards agent-based development paradigms, these models form the cornerstone of next-generation software development lifecycles. However, current benchmarks for evaluating multilingual code generation capabilities are predominantly English-centric, limiting their applicability across the global developer community. To address this limitation, we present IndicEval-XL, a comprehensive benchmark for code generation that incorporates 6 major Indic languages, collectively spoken by approximately 14\% of the world's population. Our benchmark bridges these languages with 12 programming languages, creating a robust evaluation framework. This work is particularly significant given India's representation of one-eighth of the global population and the crucial role Indic languages play in Indian society. IndicEval-XL represents a significant step toward expanding the linguistic diversity in code generation systems and evaluation frameworks. By developing resources that support multiple languages, we aim to make AI-powered development tools more inclusive and accessible to developers of various linguistic backgrounds. To facilitate further research and development in this direction, we make our dataset and evaluation benchmark publicly available at https://github.com/telekom/IndicEval-XL
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。