测试大模型翻译数字能力,发现错误率高达20%。
Investigating Numerical Translation with Large Language Models
- 构建真实业务数据的中英数字翻译数据集
- 多数开源模型在百万亿级单位翻译中错误率超20%
- 提出三种降低大单位翻译错误的策略
数字翻译不准确可能引发严重安全问题,从财务损失到医疗误判。尽管大语言模型(LLMs)在机器翻译方面取得显著进展,但其处理数字的能力尚未被充分探索。本研究系统评估了当前开源大模型在数值翻译中的可靠性。我们基于真实业务数据,构建了一个涵盖十类数值翻译的中英双语数据集。实验表明,数值翻译错误普遍存在,多数开源模型在测试场景中表现不佳。尤其在涉及百万、十亿及中文“亿”等大单位时,即使是最新版Llama3.1 8b模型,错误率也高达20%。最后,我们提出了三种缓解大单位翻译错误的潜在策略。
原文摘要 · Abstract (English)
The inaccurate translation of numbers can lead to significant security issues, ranging from financial setbacks to medical inaccuracies. While large language models (LLMs) have made significant advancements in machine translation, their capacity for translating numbers has not been thoroughly explored. This study focuses on evaluating the reliability of LLM-based machine translation systems when handling numerical data. In order to systematically test the numerical translation capabilities of currently open source LLMs, we have constructed a numerical translation dataset between Chinese and English based on real business data, encompassing ten types of numerical translation. Experiments on the dataset indicate that errors in numerical translation are a common issue, with most open-source LLMs faltering when faced with our test scenarios. Especially when it comes to numerical types involving large units like ``million", ``billion", and "yi", even the latest llama3.1 8b model can have error rates as high as 20%. Finally, we introduce three potential strategies to mitigate the numerical mistranslations for large units.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。