arXiv:2503.18250cs.CL2025-03

用短语对齐数据提升韩语迁移学习效率

PAD: Towards Efficient Data Generation for Transfer Learning Using Phrase Alignment

  • 利用标准化机器翻译生成短语对齐数据
  • 显著提升韩语模型性能,弥补传统方法不足
  • 适合资源匮乏语言的低成本迁移学习

迁移学习利用丰富的英语数据来缓解韩语等非英语语言的数据稀缺问题。本文探索了来自标准化统计机器翻译(SMT)的短语对齐数据(PAD)在提升迁移学习效率方面的潜力。通过大量实验,我们发现PAD能有效结合韩语的句法特征,克服SMT的弱点,显著提升模型性能。此外,PAD可与传统数据构建方法互补,增强其效果。该创新方法不仅提升了模型表现,还为资源匮乏语言提供了成本高效的解决方案。

原文摘要 · Abstract (English)

Transfer learning leverages the abundance of English data to address the scarcity of resources in modeling non-English languages, such as Korean. In this study, we explore the potential of Phrase Aligned Data (PAD) from standardized Statistical Machine Translation (SMT) to enhance the efficiency of transfer learning. Through extensive experiments, we demonstrate that PAD synergizes effectively with the syntactic characteristics of the Korean language, mitigating the weaknesses of SMT and significantly improving model performance. Moreover, we reveal that PAD complements traditional data construction methods and enhances their effectiveness when combined. This innovative approach not only boosts model performance but also suggests a cost-efficient solution for resource-scarce languages.

迁移学习数据生成韩语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。