让翻译和质量评估模型适应专业领域,靠精选数据与智能适配。
Toward domain-specific machine translation and quality estimation systems
- 用相似性筛选小规模专业数据,效果优于大而全的通用数据集。
- 分阶段训练提升跨语言、零样本场景下的质量评估性能。
- 结合质量评估选例,无需更新参数即可改进大模型翻译效果。
机器翻译(MT)和质量评估(QE)在通用领域表现良好,但在领域不匹配时性能下降。本文通过一系列以数据为核心的贡献,研究如何将MT和QE系统适配至专业领域。第二章提出基于相似性的数据选择方法,小规模、聚焦领域的数据集在翻译质量上超越大规模通用数据集,且计算成本更低。第三章设计分阶段的QE训练流程,融合领域自适应与轻量数据增强,在多语言、多资源设置下均提升性能,包括零样本和跨语言情形。第四章研究子词分词与词汇表在微调中的作用,对齐的分词-词汇配置带来稳定训练与更优翻译质量,而配置不匹配则降低性能。第五章提出一种基于QE引导的上下文学习方法,仅通过选择高质量示例即可提升大语言模型翻译效果,无需参数更新,且支持无参考设置,减少对单一参考集依赖。结果表明,领域适配依赖于数据选择、表示方式与高效适配策略。本文为构建在特定领域中可靠运行的MT与QE系统提供了有效方法。
原文摘要 · Abstract (English)
Machine Translation (MT) and Quality Estimation (QE) perform well in general domains but degrade under domain mismatch. This dissertation studies how to adapt MT and QE systems to specialized domains through a set of data-focused contributions. Chapter 2 presents a similarity-based data selection method for MT. Small, targeted in-domain subsets outperform much larger generic datasets and reach strong translation quality at lower computational cost. Chapter 3 introduces a staged QE training pipeline that combines domain adaptation with lightweight data augmentation. The method improves performance across domains, languages, and resource settings, including zero-shot and cross-lingual cases. Chapter 4 studies the role of subword tokenization and vocabulary in fine-tuning. Aligned tokenization-vocabulary setups lead to stable training and better translation quality, while mismatched configurations reduce performance. Chapter 5 proposes a QE-guided in-context learning method for large language models. QE models select examples that improve translation quality without parameter updates and outperform standard retrieval methods. The approach also supports a reference-free setup, reducing reliance on a single reference set. These results show that domain adaptation depends on data selection, representation, and efficient adaptation strategies. The dissertation provides methods for building MT and QE systems that perform reliably in domain-specific settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。