arXiv:2511.07461cs.CLcs.AI2025-11中稿 · WMT 2025被引 3

用双阶段方法提升术语翻译准确率,效果优于传统强制约束。

It Takes Two: A Dual Stage Approach for Terminology-Aware Translation

  • 先用合成数据微调术语感知模型,再用提示词驱动大模型后处理
  • 在WMT 2025术语任务上表现更优,术语遵循更灵活自然
  • 适合需要高精度术语翻译的领域,如法律、医疗文档

本文提出DuTerm,一种面向术语约束机器翻译的双阶段架构。系统结合术语感知NMT模型(通过大规模合成数据微调)与基于提示词的LLM后编辑模块。LLM阶段对NMT输出进行优化并确保术语一致性。我们在英语→德语、英语→西班牙语、英语→俄语任务上,使用WMT 2025术语共享任务语料库进行评估。结果表明,由LLM实现的灵活、上下文驱动的术语处理方式,持续优于严格约束策略。研究揭示关键权衡:LLM在高质量翻译中作为上下文驱动的修正者表现更佳,而非生成者。

原文摘要 · Abstract (English)

This paper introduces DuTerm, a novel two-stage architecture for terminology-constrained machine translation. Our system combines a terminology-aware NMT model, adapted via fine-tuning on large-scale synthetic data, with a prompt-based LLM for post-editing. The LLM stage refines NMT output and enforces terminology adherence. We evaluate DuTerm on English-to German, English-to-Spanish, and English-to-Russian with the WMT 2025 Terminology Shared Task corpus. We demonstrate that flexible, context-driven terminology handling by the LLM consistently yields higher quality translations than strict constraint enforcement. Our results highlight a critical trade-off, revealing that an LLM's work best for high-quality translation as context-driven mutators rather than generators.

术语翻译双阶段大模型后处理NMT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。