用吉姆玛3模型打造卢森堡语翻译系统,效果显著优于基线。
LuxMT Technical Report
- 基于吉姆玛3 27B微调,专攻卢森堡语转法语、英语。
- 在卢森堡语→法语/英语上表现优于基线,甚至对德语也有效。
- 提出嵌入向量筛选方法,可辅助质量评估但需谨慎使用。
我们提出LuxMT,一个基于Gemma 3 27B的机器翻译系统,专门用于将卢森堡语(LB)翻译为法语(FR)和英语(EN)。为评估性能,构建了包含LB-FR、LB-EN及回译数据的新基准,数据源自卢森堡旅游杂志Luci的人工翻译内容。训练数据来自多语言卢森堡新闻语料库LuxAlign,以及经Google Translate增强的卢森堡议会记录。通过使用LuxEmbedder(LB句向量)过滤低等价句子对,提升数据质量。实验表明,LuxMT在多个方向上均显著优于Gemma 3基线,即使训练未含德语(DE),对卢森堡语→德语翻译也有明显提升。此外,探索发现LuxEmbedder与主流参考指标有强相关性,具备潜在质量估计价值,但其有效性仍需进一步验证,建议谨慎使用。
原文摘要 · Abstract (English)
We introduce LuxMT, a machine translation system based on Gemma 3 27B and fine-tuned for translation from Luxembourgish (LB) into French (FR) and English (EN). To assess translation performance, we construct a novel benchmark covering LB-FR, LB-EN, and LB-FR using human-translated data from Luci, a tourist magazine about Luxembourg. Training data stems from LuxAlign, a parallel corpus of multilingual Luxembourgish news articles, and LB parliamentary transcripts augmented with Google Translate. We filter the data using LuxEmbedder, LB sentence embeddings, to remove low-equivalence segment-pairs. Overall, LuxMT's results suggest strong improvements over the Gemma 3 baseline, even for translating LB to German (DE), despite the training data not containing any DE. We also explore LuxEmbedder's potential to be used as a quality estimation metric and find strong correlations with other reference-based metrics. However, we call for further research to fully assess the metric's utility and advise using it with caution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。