arXiv:2503.09454cs.CL2025-03EMNLP被引 6

大模型能通过语法书学习新语言,但复杂度越高效果越差。

Explicit Learning and the LLM in Machine Translation

  • 用语法书解释让大模型学习构造语言
  • 模型在简单语言上表现良好,复杂度上升时性能下降
  • 适合研究低资源语言的可学习性与训练策略

本研究探讨大语言模型(LLM)通过语法书中的解释来学习新语言的能力,称为“显式学习”。为严格评估该能力,设计了英语与由拉丁语或法语经特定加密方式生成的构造语言之间的受控翻译实验。结果表明,尽管先前研究持怀疑态度,但LLM确实具备可测量的显式学习能力。然而,随着需学习语言现象的复杂度增加,该能力显著减弱。在自定义思维链上进行监督微调能显著提升模型表现,但难以泛化至类型学上新颖或更复杂的语言特征。研究提示需构建更丰富的训练数据集,并探索替代微调策略,以进一步提升大模型的显式学习能力,从而惠及那些仅存于语法书中、缺乏大规模语料的低资源语言。

原文摘要 · Abstract (English)

This study explores an LLM's ability to learn new languages using explanations found in a grammar book, a process we term "explicit learning." To rigorously assess this ability, we design controlled translation experiments between English and constructed languages generated, through specific cryptographic means, from Latin or French. Contrary to previous studies, our results demonstrate that LLMs do possess a measurable capacity for explicit learning. This ability, however, diminishes as the complexity of the linguistic phenomena to be learned increases. Supervised fine-tuning on ad hoc chains of thought significantly enhances LLM performance but struggles to generalize to typologically novel or more complex linguistic features. These findings point to the need for more diverse training sets and alternative fine-tuning strategies to further improve explicit learning by LLMs, benefiting low-resource languages typically described in grammar books but lacking extensive corpora.

大模型机器翻译显式学习低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。