用大模型自动简化输入文本,能有效提升机器翻译质量。
Automatic Input Rewriting Improves Translation with Large Language Models
- 用21种方法自动重写输入文本,以提升翻译效果。
- 文本简化是最有效的策略,结合可译性评估可进一步提升性能。
- 简化后内容与原文语义高度一致,适合追求高质量翻译的场景。
我们通过实证研究评估了21种自动输入重写方法,使用3个开源大模型将英语翻译成6种目标语言。结果表明,文本简化是效果最佳的通用重写策略,结合可译性质量评估可进一步优化。人工评估确认,简化后的输入及其生成的机器翻译在很大程度上保留了原文语义。这些发现表明,大模型辅助的输入重写是提升翻译质量的有前景方向。
原文摘要 · Abstract (English)
Can we improve machine translation (MT) with LLMs by rewriting their inputs automatically? Users commonly rely on the intuition that well-written text is easier to translate when using off-the-shelf MT systems. LLMs can rewrite text in many ways but in the context of MT, these capabilities have been primarily exploited to rewrite outputs via post-editing. We present an empirical study of 21 input rewriting methods with 3 open-weight LLMs for translating from English into 6 target languages. We show that text simplification is the most effective MT-agnostic rewrite strategy and that it can be improved further when using quality estimation to assess translatability. Human evaluation further confirms that simplified rewrites and their MT outputs both largely preserve the original meaning of the source and MT. These results suggest LLM-assisted input rewriting as a promising direction for improving translations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。