选对数据比调模型更能提升大语言模型翻译质量
Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis
- 用语义相似度筛选训练数据,效果优于词频或几何方法
- 数据差异小于3%时,翻译性能仍有明显差别
- 适合想用少量数据高效优化翻译模型的研究者
我们研究了数据选择对开放大语言模型机器翻译微调的影响。以日英语料库为基础,在受控训练条件下对比了五种数据选择器:TF-IDF、COMET Kiwi、QuRate、FD-Score 和随机选择。结果表明,语义类选择器始终优于基于词频和几何特征的启发式方法;即使所选数据差异不足3%,模型性能仍出现显著变化,凸显微调对数据质量的高度敏感性。
原文摘要 · Abstract (English)
We investigated the impact of data selection on machine translation fine-tuning for open LLMs. Using Japanese-English corpora, we compare five selectors: TF-IDF, COMET Kiwi, QuRate, FD-Score, and random selection, under controlled training conditions. We observed that semantic selectors consistently outperform lexical and geometry-based heuristics, and that even when the selected data differ by less than 3%, the impact on model performance is substantial, underscoring the sensitivity of fine-tuning to data quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。