评估商用机器翻译在疫情低资源语言中的实用性
Is MT Ready for the Next Crisis or Pandemic?
- 用TICO-19数据集测试四种商用MT系统
- 低资源语言翻译准确率普遍低于60%
- 适合公共卫生与国际救援人员参考
危机时刻的沟通至关重要。然而,政府、援助机构、医生与受助人群之间常存在语言障碍。商用机器翻译系统可作为应对工具,但其在低资源语言、尤其疫情或医疗场景下的表现如何?本研究使用包含高优先级语言的TICO-19数据集,评估四种商用机器翻译系统,并基于输出译文的可用性,评估当前对下一次疫情或大流行病的准备程度。
原文摘要 · Abstract (English)
Communication in times of crisis is essential. However, there is often a mismatch between the language of governments, aid providers, doctors, and those to whom they are providing aid. Commercial MT systems are reasonable tools to turn to in these scenarios. But how effective are these tools for translating to and from low resource languages, particularly in the crisis or medical domain? In this study, we evaluate four commercial MT systems using the TICO-19 dataset, which is composed of pandemic-related sentences from a large set of high priority languages spoken by communities most likely to be affected adversely in the next pandemic. We then assess the current degree of ``readiness'' for another pandemic (or epidemic) based on the usability of the output translations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。