用大模型迭代生成难译文本,提升机器翻译评测的区分度。
Generating Difficult-to-Translate Texts
- 通过大模型不断优化文本,基于目标翻译模型反馈生成更难译例。
- 生成的测试例对主流模型挑战性更强,且保持自然文本多样性。
- 适合用于评估和发现翻译模型缺陷,尤其对新模型测试有效。
真实世界中的机器翻译评测基准正迅速过时,因为多数样本对当前先进翻译模型而言过于简单,无法有效区分模型优劣或暴露其弱点。现有生成难译样本的方法,如子采样或从头合成,要么难以识别真正难点,要么缺乏多样性和自然性。受人类专家反复试探模型漏洞的启发,我们提出 MT-breaker:一种利用大语言模型迭代优化源文本以提升翻译难度的方法。该方法通过持续向目标机器翻译模型提问,指导生成更具挑战性的测试例。生成的文本在针对特定翻译模型时更具难度,但这种难度也具备跨模型和跨语言的迁移性,同时保持了自然文本的多样性。
原文摘要 · Abstract (English)
Machine translation benchmarks sourced from the real world are quickly obsoleted, due to most examples being easy for state-of-the-art translation models. This limits the benchmark's ability to distinguish which model is better or to reveal models' weaknesses. Current methods for creating difficult test cases, such as subsampling or from-scratch synthesis, either fall short of identifying difficult examples or suffer from a lack of diversity and naturalness. Inspired by the iterative process of human experts probing for model failures, we propose MT-breaker, a method where a large language model iteratively refines a source text to increase its translation difficulty. The LLM iteratively queries a target machine translation model to guide its generation of difficult examples. Our approach generates examples that are more challenging for the target MT model while preserving the diversity of natural texts. While the examples are tailored to a particular machine translation model during the generation, the difficulty also transfers to other models and languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。