评测本地部署大模型对波罗的海三语的支持能力
Localizing AI: Evaluating Open-Weight Language Models for Languages of Baltic States
- 测试多个开源大模型在立陶宛语等小语种上的表现
- 部分模型翻译接近商用水平,但每20词就有1词错误
- 适合关注数据隐私的政府与国防机构参考
尽管大型语言模型(LLMs)已改变现代语言技术的预期,但数据隐私问题常限制在欧盟境外托管的商用LLM使用,影响政府、国防等敏感领域的应用。本文评估本地可部署的开源权重模型对立陶宛语、拉脱维亚语和爱沙尼亚语等小语种的支持程度。测试了Llama~3、Gemma~2、Phi和NeMo等主流多语言开源模型在机器翻译、多项选择问答和自由文本生成任务上的表现。结果表明,部分模型如Gemma~2的翻译性能接近顶尖商用模型,但多数模型仍表现不佳。最令人意外的是,尽管翻译表现接近前沿水平,所有开源多语言大模型在词汇层面均存在幻觉,错误率至少为每20个词出现1次。
原文摘要 · Abstract (English)
Although large language models (LLMs) have transformed our expectations of modern language technologies, concerns over data privacy often restrict the use of commercially available LLMs hosted outside of EU jurisdictions. This limits their application in governmental, defence, and other data-sensitive sectors. In this work, we evaluate the extent to which locally deployable open-weight LLMs support lesser-spoken languages such as Lithuanian, Latvian, and Estonian. We examine various size and precision variants of the top-performing multilingual open-weight models, Llama~3, Gemma~2, Phi, and NeMo, on machine translation, multiple-choice question answering, and free-form text generation. The results indicate that while certain models like Gemma~2 perform close to the top commercially available models, many LLMs struggle with these languages. Most surprisingly, however, we find that these models, while showing close to state-of-the-art translation performance, are still prone to lexical hallucinations with errors in at least 1 in 20 words for all open-weight multilingual LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。