arXiv:2410.17145cs.CLcs.AI2024-10中稿 · EMNLP被引 2

大模型在泰语翻译中受限于算力,专用模型表现更优。

Can General-Purpose Large Language Models Generalize to English-Thai Machine Translation ?

  • 测试大模型与专用模型在英泰翻译任务中的表现
  • 4比特量化下大模型翻译失效,专用模型仍稳定
  • 适合资源受限场景的翻译系统设计

大语言模型(LLMs)在常见任务上表现良好,但在低资源和低算力环境下泛化能力差。我们通过在英语-泰语机器翻译和代码混用数据集上测试多种LLMs和专用翻译模型,发现当计算约束更严格时,如采用4比特量化,LLMs无法有效翻译。相比之下,专用模型在计算需求相当或更低的情况下,始终优于LLMs。这凸显了在资源受限条件下,专用模型对性能保持的重要性。

原文摘要 · Abstract (English)

Large language models (LLMs) perform well on common tasks but struggle with generalization in low-resource and low-computation settings. We examine this limitation by testing various LLMs and specialized translation models on English-Thai machine translation and code-switching datasets. Our findings reveal that under more strict computational constraints, such as 4-bit quantization, LLMs fail to translate effectively. In contrast, specialized models, with comparable or lower computational requirements, consistently outperform LLMs. This underscores the importance of specialized models for maintaining performance under resource constraints.

机器翻译大模型量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。