arXiv:2503.20959cs.CLcs.AI2025-03被引 8

探讨机器翻译的副作用与应对策略,关注碳排放与译者影响。

Sociotechnical Effects of Machine Translation

  • 用小模型和微调替代从头训练,降低碳足迹。
  • 大型模型训练耗能高,单次训练可产生数公斤二氧化碳。
  • 适合关注AI伦理、可持续计算与紧急救援应用的研究者。

尽管机器翻译(MT)具有实用价值,但其伴随的副作用和风险不容忽视。随着神经机器翻译及大语言模型(LLMs)的发展,跨国企业构建的模型规模巨大,训练成本高昂,耗电量大,单次训练可释放数公斤二氧化碳。相比之下,高性能的小型模型碳足迹显著更低,微调预训练模型也避免了从头训练的需求。本文还探讨了机器翻译对译者及其他用户可能造成的负面影响,涉及数据版权与所有权问题,以及数据使用中的伦理考量。最后,文章展示了在危机场景中合理使用机器翻译可挽救生命,并提出具体实施方法。

原文摘要 · Abstract (English)

While the previous chapters have shown how machine translation (MT) can be useful, in this chapter we discuss some of the side-effects and risks that are associated, and how they might be mitigated. With the move to neural MT and approaches using Large Language Models (LLMs), there is an associated impact on climate change, as the models built by multinational corporations are massive. They are hugely expensive to train, consume large amounts of electricity, and output huge volumes of kgCO2 to boot. However, smaller models which still perform to a high level of quality can be built with much lower carbon footprints, and tuning pre-trained models saves on the requirement to train from scratch. We also discuss the possible detrimental effects of MT on translators and other users. The topics of copyright and ownership of data are discussed, as well as ethical considerations on data and MT use. Finally, we show how if done properly, using MT in crisis scenarios can save lives, and we provide a method of how this might be done.

机器翻译碳足迹伦理问题危机响应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。