对比大模型与小模型在马耳他语上的表现,发现小模型更优。
MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP
- 用11项任务测试55个大模型,评估其在低资源语言马耳他语的表现。
- 小规模微调模型在所有任务中均优于多数大模型,尤其生成任务表现更佳。
- 预训练时接触过马耳他语是性能关键,微调比提示工程性价比更高。
大型语言模型(LLMs)在多种自然语言处理任务中表现出色,主要因其泛化能力及无需额外训练即可执行任务的特性。然而,它们在低资源语言上的效果仍有限。本研究使用新提出的涵盖11项判别与生成任务的基准,评估了55个公开可用的LLM在马耳他语上的表现。实验显示,许多模型表现不佳,尤其在生成任务中;而较小的微调模型在所有任务中表现更优。多维度分析揭示,预训练阶段是否包含马耳他语以及指令微调是影响性能的最关键因素。我们还探讨了微调与提示之间的权衡:尽管微调初始成本较高,但其性能更好且推理成本更低。本文旨在呼吁更具包容性的语言技术发展,建议低资源语言研究者考虑采用更传统的语言建模范式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable performance across various Natural Language Processing (NLP) tasks, largely due to their generalisability and ability to perform tasks without additional training. However, their effectiveness for low-resource languages remains limited. In this study, we evaluate the performance of 55 publicly available LLMs on Maltese, a low-resource language, using a newly introduced benchmark covering 11 discriminative and generative tasks. Our experiments highlight that many models perform poorly, particularly on generative tasks, and that smaller fine-tuned models often perform better across all tasks. From our multidimensional analysis, we investigate various factors impacting performance. We conclude that prior exposure to Maltese during pre-training and instruction-tuning emerges as the most important factor. We also examine the trade-offs between fine-tuning and prompting, highlighting that while fine-tuning requires a higher initial cost, it yields better performance and lower inference costs. Through this work, we aim to highlight the need for more inclusive language technologies and recommend that researchers working with low-resource languages consider more "traditional" language modelling approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。