首次为646种编程语言建立资源分级体系,揭示代码生成数据严重不均。
A Taxonomy of Programming Languages for Code Generation
- 按资源丰富度将编程语言分为四类,基于7个主流语料库数据
- 仅1.9%高资源语言贡献74.6%代码令牌,71.7%低资源语言仅占1.0%
- 为多语言大模型评估与数据集构建提供系统性依据
全球7000多种语言在自然语言处理资源上差异显著,促使研究者对它们进行资源丰富度分类(Joshi等,2020)。编程语言同样存在类似差距,但尚未建立代码领域的资源分级体系。随着大语言模型(LLMs)在代码生成能力上的提升,此类分类变得至关重要。为此,我们提出首个可复现的编程语言资源分类方法,将646种语言划分为四个层级。结果显示,仅1.9%的高资源语言(Tier 3)贡献了七个主要语料库中74.6%的代码令牌,而71.7%的低资源语言(Tier 0)仅贡献1.0%。对层内不平等、分布离散性和偏态的统计分析表明,这种失衡既极端又系统化。本研究为数据集构建和多语言大模型的分层评估提供了原则性框架。
原文摘要 · Abstract (English)
The world's 7,000+ languages vary widely in the availability of resources for NLP, motivating efforts to systematically categorize them by their degree of resourcefulness (Joshi et al., 2020). A similar disparity exists among programming languages (PLs); however, no resource-tier taxonomy has been established for code. As large language models (LLMs) grow increasingly capable of generating code, such a taxonomy becomes essential. To fill this gap, we present the first reproducible PL resource classification, grouping 646 languages into four tiers. We show that only 1.9% of languages (Tier 3, High) account for 74.6% of all tokens in seven major corpora, while 71.7% of languages (Tier 0, Scarce) contribute just 1.0%. Statistical analyses of within-tier inequality, dispersion, and distributional skew confirm that this imbalance is both extreme and systematic. Our results provide a principled framework for dataset curation and tier-aware evaluation of multilingual LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。