用少量地道黎巴嫩方言数据微调,效果优于大量非本地数据。
Fine-Tuning LLMs for Low-Resource Dialect Translation: The Case of Lebanese
- 用文化真实的小规模方言数据微调模型,比大而无趣的数据更有效。
- 对比式微调+提示显著提升翻译质量,错误样本反而有帮助。
- 适合关注低资源方言翻译与文化适配的研究者或开发者。
本文研究大型语言模型在低资源黎巴嫩方言翻译中的表现,重点对比文化真实数据与更大规模翻译数据的效果。采用开源Aya23模型,比较基础、对比式和语法提示三种微调方法。实验表明,使用小规模但具有文化意识的黎巴嫩语数据集(LW)微调的模型,始终优于基于大规模非本地数据训练的模型。最佳结果来自对比式微调结合对比提示,说明暴露于错误示例有益。为实现真实评估,我们构建了源自本土内容的新基准LebEval,并与现有FLoRes基准对比。研究挑战了“数据越多越好”的传统观念,强调文化真实性在方言翻译中的关键作用。相关数据集与代码已公开于Github。
原文摘要 · Abstract (English)
This paper examines the effectiveness of Large Language Models (LLMs) in translating the low-resource Lebanese dialect, focusing on the impact of culturally authentic data versus larger translated datasets. We compare three fine-tuning approaches: Basic, contrastive, and grammar-hint tuning, using open-source Aya23 models. Experiments reveal that models fine-tuned on a smaller but culturally aware Lebanese dataset (LW) consistently outperform those trained on larger, non-native data. The best results were achieved through contrastive fine-tuning paired with contrastive prompting, which indicates the benefits of exposing translation models to bad examples. In addition, to ensure authentic evaluation, we introduce LebEval, a new benchmark derived from native Lebanese content, and compare it to the existing FLoRes benchmark. Our findings challenge the "More Data is Better" paradigm and emphasize the crucial role of cultural authenticity in dialectal translation. We made our datasets and code available on Github.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。