arXiv:2602.14675cs.CL2026-02

构建濒危语言皮埃蒙特语数据集,测试大模型在非标准拼写的处理能力

Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthography

  • 由母语者以自然拼写方式翻译意大利语-皮埃蒙特语平行句
  • 大模型在主题分类上接近意法英表现,但皮埃蒙特语分词性能较差
  • 适合研究低资源语言、非标准拼写及多语言模型泛化能力的学者

我们提出一个众包构建的皮埃蒙特语数据集,该语言是意大利西北部的濒危罗曼语。数据集包含145对意大利语-皮埃蒙特语平行句子,源自Flores+,由母语者以自然拼写风格翻译,而非遵循标准化拼写规范,并附有人工词对齐。我们利用该资源对多个大语言模型进行分词一致性、主题分类和机器翻译的基准测试。分析显示,相较于高资源罗曼语,皮埃蒙特语存在分词惩罚现象,但大模型的主题分类性能已接近意大利语、法语和英语水平。机器翻译结果呈不对称性:从皮埃蒙特语到高资源语言的翻译效果尚可,但从高资源语言生成皮埃蒙特语仍具挑战性。数据集与代码已公开。

原文摘要 · Abstract (English)

We present a crowdsourced dataset for Piedmontese, an endangered Romance language of northwestern Italy. The dataset comprises 145 Italian-Piedmontese parallel sentences derived from Flores+, with translations produced by speakers writing in their natural orthographic style rather than adhering to standardized conventions, along with manual word alignment. We use this resource to benchmark several large language models on tokenization parity, topic classification, and machine translation. Our analysis reveals that Piedmontese incurs a tokenization penalty relative to higher-resource Romance languages, yet LLMs achieve classification performance approaching that of Italian, French, and English. Machine translation results are asymmetric: models translate adequately from Piedmontese into high-resource languages, but generation into Piedmontese remains challenging. The dataset and code are publicly released.

低资源语言非标准拼写大模型评测皮埃蒙特语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。