测试24个开源模型在僧伽罗语不同书写形式下的表现,发现罗马字母文本性能下降超300倍。
Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala
- 用困惑度评估24个模型在僧伽罗语的纯拉丁、纯Unicode和混合书写形式下的表现。
- 从Unicode到罗马化文本,平均性能下降超过300倍,小模型反而常优于大模型。
- 单一脚本评估严重低估实际应用挑战,适合多语言低资源场景的模型选型参考。
语言模型(LMs)在低资源、形态丰富的语言如僧伽罗语上的表现仍鲜有研究,尤其在数字交流中脚本变异的问题。僧伽罗语具有书写双重性:正式场合使用Unicode,社交媒体则以罗马化文本为主,实际中混合脚本极为常见。本文基于多种文本来源,对24个开源语言模型在Unicode、罗马化及混合脚本的僧伽罗语上进行困惑度评估。结果表明,模型存在显著脚本敏感性,从Unicode到罗马化文本,中位性能下降超过300倍。关键的是,模型规模与脚本处理能力无相关性——较小模型常优于大出28倍的架构。此外,Unicode性能能较好预测混合脚本鲁棒性,但无法预测罗马化能力。这些发现确立了僧伽罗语语言模型的基准能力,并为多脚本低资源环境中的模型选择提供了实践指导。
原文摘要 · Abstract (English)
The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly regarding script variation in digital communication. Sinhala exhibits script duality, with Unicode used in formal contexts and Romanized text dominating social media, while mixed-script usage is common in practice. This paper benchmarks 24 open-source LMs on Unicode, Romanized and mixed-script Sinhala using perplexity evaluation across diverse text sources. Results reveal substantial script sensitivity, with median performance degradation exceeding 300 times from Unicode to Romanized text. Critically, model size shows no correlation with script-handling competence, as smaller models often outperform architectures 28 times larger. Unicode performance strongly predicts mixed-script robustness but not Romanized capability, demonstrating that single-script evaluation substantially underestimates real-world deployment challenges. These findings establish baseline LM capabilities for Sinhala and provide practical guidance for model selection in multi-script low-resource environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。