arXiv:2507.09509cs.CL2025-07Conference of the …被引 2

研究英文提示错误对机器翻译模型的影响,发现错误类型差异大且模型仍可应对乱码

How Important is `Perfect' English for Machine Translation Prompts?

  • 系统测试人类可读和合成错误对翻译提示的影响
  • 字符级噪声比短语级干扰更严重,错误越多性能越差
  • 模型误判多因指令理解偏差,非直接翻译质量下降

大型语言模型(LLMs)在近期机器翻译评测中表现优异,但对提示词中的错误和扰动敏感。我们系统评估了用户提示中的人类可接受错误与合成错误对两个相关任务的影响:机器翻译与机器翻译评估。通过定量分析与定性洞察,揭示了提示质量对翻译性能的显著影响:错误较多时,即使优质提示也可能劣于无错但基础的提示。不同噪声类型影响各异,字符级及组合型噪声对性能的损害超过短语级扰动。定性分析显示,低质量提示主要导致指令遵循能力下降,而非直接影响翻译质量本身。此外,即便提示充斥随机噪声而人类无法辨识,模型仍能完成翻译。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved top results in recent machine translation evaluations, but they are also known to be sensitive to errors and perturbations in their prompts. We systematically evaluate how both humanly plausible and synthetic errors in user prompts affect LLMs' performance on two related tasks: Machine translation and machine translation evaluation. We provide both a quantitative analysis and qualitative insights into how the models respond to increasing noise in the user prompt. The prompt quality strongly affects the translation performance: With many errors, even a good prompt can underperform a minimal or poor prompt without errors. However, different noise types impact translation quality differently, with character-level and combined noisers degrading performance more than phrasal perturbations. Qualitative analysis reveals that lower prompt quality largely leads to poorer instruction following, rather than directly affecting translation quality itself. Further, LLMs can still translate in scenarios with overwhelming random noise that would make the prompt illegible to humans.

机器翻译提示工程语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。