用大模型生成高质量数据,提升越日翻译效果
VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation
- 用思维链提示的大模型生成合成数据,解决语料稀缺问题
- 在1.8亿参数模型上用QLoRA高效微调,实现高性能翻译
- 适合低资源语言对研究者和实际应用开发者参考
基于Transformer的神经机器翻译虽有显著进展,但在越南语-日语(Vi-Ja)等低资源语言对上仍面临挑战,如平行语料稀疏及语言文化差异。近期大型语言模型(LLMs)结合强化学习,在强推理能力下可生成高质量合成数据。本文提出VNJPTranslate,一个系统性解决Vi-Ja翻译的流水线:首先通过语料分析识别难点片段,利用带思维链提示的先进大模型进行针对性数据增强;随后在具备能力但参数量少的自回归模型(具体为基于Qwen架构、经微调的1.8亿参数Sailor模型)上,采用高效的微调技术(Unsloth + QLoRA),构建出实用且性能优异的翻译系统。该综合方法显著优于现有基线。
原文摘要 · Abstract (English)
Neural Machine Translation (NMT) driven by Transformer architectures has advanced significantly, yet faces challenges with low-resource language pairs like Vietnamese-Japanese (Vi-Ja). Issues include sparse parallel data and handling linguistic/cultural nuances. Recent progress in Large Language Models (LLMs) with strong reasoning, often refined via Reinforcement Learning (RL), enables high-quality synthetic data generation. We introduce VNJPTranslate, a pipeline designed to systematically address the Vi-Ja translation task. It features a targeted data augmentation strategy using advanced LLMs with Chain-of-Thought prompting for challenging segments identified via corpus analysis. Subsequently, we employ efficient fine-tuning techniques (Unsloth with QLoRA) on a capable, low-parameter autoregressive model (specifically, a fine-tuned version of the 1.8B parameter Sailor model, which is based on the Qwen architecture) to create a practical and high-performing translation system. This integrated approach aims to improve Vi-Ja translation quality significantly over existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。