arXiv:2601.09725cs.CL2026-01

测试并提升英马拉地语翻译对标点错误的鲁棒性

Assessing and Improving Punctuation Robustness in English-Marathi Machine Translation

  • 构建54组带歧义标点的英马拉地语对照句,用于诊断模型缺陷
  • 修复标点后再翻译或直接微调,均显著提升翻译准确率
  • 大模型在该任务上表现不如专用方法,提示需专门研究

神经机器翻译系统高度依赖显式标点来消除源句语义歧义。输入用户生成的、可能缺失或错误标点的句子,会导致流畅但语义灾难性的翻译结果。本文针对英马拉地语翻译,提出并解决标点鲁棒性问题。首先,我们构建了名为「Viram」的人工标注诊断基准,包含54组标点歧义的英马拉地语句子对,用以压力测试现有NMT系统。其次,评估两种简单修复策略:级联式「先恢复再翻译」与直接微调。实验结果与分析表明,两种策略均带来显著的翻译性能提升。此外,我们发现当前大型语言模型在处理此类句子时,鲁棒性反而低于这些任务专用策略,凸显该领域仍需深入研究。代码与数据集已开源:https://github.com/KaustubhShejole/Viram_Marathi。

原文摘要 · Abstract (English)

Neural Machine Translation (NMT) systems rely heavily on explicit punctuation cues to resolve semantic ambiguities in a source sentence. Inputting user-generated sentences, which are likely to contain missing or incorrect punctuation, results in fluent but semantically disastrous translations. This work attempts to highlight and address the problem of punctuation robustness of NMT systems through an English-to-Marathi translation. First, we introduce \textbf{\textit{Viram}}, a human-curated diagnostic benchmark of 54 punctuation-ambiguous English-Marathi sentence pairs to stress-test existing NMT systems. Second, we evaluate two simple remediation strategies: cascade-based \textit{restore-then-translate} and \textit{direct fine-tuning}. Our experimental results and analysis demonstrate that both strategies yield substantial NMT performance improvements. Furthermore, we find that current Large Language Models (LLMs) exhibit relatively poorer robustness in translating such sentences than these task-specific strategies, thus necessitating further research in this area. The code and dataset are available at https://github.com/KaustubhShejole/Viram_Marathi.

机器翻译标点鲁棒性多语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。