用大模型辅助人工校对,高效生成高质量机器翻译语料。
Efficient Machine Translation Corpus Generation: Integrating Human-in-the-Loop Post-Editing with Large Language Models
- 大模型生成初译,人工校对环节嵌入智能推荐与质量评估。
- 降低人工标注负担,提升翻译质量与生成效率。
- 适合需要快速构建高质量翻译数据集的研究者与团队。
本文提出一种高效的机器翻译语料生成方法,将半自动化的人工介入后编辑与大型语言模型(LLMs)相结合,以提升效率与翻译质量。在已有实时训练自定义机器翻译质量评估指标的基础上,系统引入了增强翻译合成与辅助注释分析等新型大模型功能,分别改进初始翻译结果和质量评估。此外,还采用大模型驱动的伪标签生成与翻译推荐系统,在特定场景下减少人工标注负担。该方法不仅保持了成本降低和后编辑质量提升的优势,还为利用前沿大模型技术开辟了新路径。项目源代码已开源,支持社区协作发展。演示视频可访问。
原文摘要 · Abstract (English)
This paper introduces an advanced methodology for machine translation (MT) corpus generation, integrating semi-automated, human-in-the-loop post-editing with large language models (LLMs) to enhance efficiency and translation quality. Building upon previous work that utilized real-time training of a custom MT quality estimation metric, this system incorporates novel LLM features such as Enhanced Translation Synthesis and Assisted Annotation Analysis, which improve initial translation hypotheses and quality assessments, respectively. Additionally, the system employs LLM-Driven Pseudo Labeling and a Translation Recommendation System to reduce human annotator workload in specific contexts. These improvements not only retain the original benefits of cost reduction and enhanced post-edit quality but also open new avenues for leveraging cutting-edge LLM advancements. The project's source code is available for community use, promoting collaborative developments in the field. The demo video can be accessed here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。