用自适应优化训练更生动的字幕翻译大模型
From Utterance to Vividity: Training Expressive Subtitle Translation LLM via Adaptive Local Preference Optimization
- 提出ALPO方法实现细粒度偏好对齐
- 在多维度评估中表现优异,提升翻译生动性
- 适合需要高质量字幕翻译的场景
大型语言模型(LLMs)的发展显著提升了机器翻译的通用能力。然而,随着应用场景日益复杂,LLMs在垂直领域翻译中的局限性逐渐显现。本文聚焦于构建满足领域定制需求的翻译大模型,以视觉媒体字幕翻译为研究对象,探索如何训练表达性强、生动自然的翻译大模型。通过分析字幕翻译及其他字面与意译类任务,验证了LLM作为奖励模型和评价器在翻译任务中的可靠性。为此,我们构建并发布了多向字幕平行语料库数据集,并提出自适应局部偏好优化(ALPO)方法,以解决细粒度偏好对齐问题。实验结果表明,ALPO在多维度翻译质量评估中表现突出。
原文摘要 · Abstract (English)
The rapid development of Large Language Models (LLMs) has significantly enhanced the general capabilities of machine translation. However, as application scenarios become more complex, the limitations of LLMs in vertical domain translations are gradually becoming apparent. In this study, we focus on how to construct translation LLMs that meet the needs of domain customization. We take visual media subtitle translation as our topic and explore how to train expressive and vivid translation LLMs. We investigated the situations of subtitle translation and other domains of literal and liberal translation, verifying the reliability of LLM as reward model and evaluator for translation. Additionally, to train an expressive translation LLM, we constructed and released a multidirectional subtitle parallel corpus dataset and proposed the Adaptive Local Preference Optimization (ALPO) method to address fine-grained preference alignment. Experimental results demonstrate that ALPO achieves outstanding performance in multidimensional evaluation of translation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。