精心设计输入可让小模型用少量数据学会语言规则
Analogical Structure, Minimal Contextual Cues and Contrastive Distractors: Input Design for Sample-Efficient Linguistic Rule Induction
- 用类比结构组织输入,提升学习效率
- 小模型用数百到千条数据即达高F1分数
- 适合研究高效语言学习与模型机制对比
大型语言模型在许多任务上表现优异,但其训练过程难以揭示哪些输入特性有助于高效的语言规则学习。本文探究三种受认知启发的输入设计原则对样本高效语言规则归纳的支持作用:类比结构、对比学习和最小上下文线索,并与大模型在相同控制任务上的表现进行对比。通过结构化句子补全任务测试英语动词交替现象,轻量级模型在数百至一千个示例上训练后即可在这些任务上获得高F1值。消融实验表明,类比组织是提升样本效率的主要因素,而对比干扰项和最小上下文进一步促进性能提升。同时在相同任务上评估零样本和少样本大模型表现。在该控制环境下,轻量级模型以远少的任务特定数据达到更高F1值。我们将此差异视为不同学习范式间的比较,而非对大模型的一般性评判。结果表明,精心设计的输入可支持语言规则的高效学习,并揭示了轻量模型与提示大模型不同的学习特征。
原文摘要 · Abstract (English)
Large language models achieve strong performance on many tasks, but their training makes it hard to see which properties of the input support efficient linguistic rule learning. We ask how three cognitively-inspired principles of input design support sample-efficient linguistic rule induction: analogical structure, contrastive learning, and minimal contextual cue. We also ask how their effects compare to those of LLMs on the same controlled tasks. We implement these principles in structured sentence completion tasks that test English verb alternations. Lightweight models trained on hundreds to one-thousand such examples learn the alternation rules with high F1 on these tasks. Ablation studies show that analogical organisation is the main driver of sample efficiency, and contrastive distractors and minimal context help further gains. We also evaluate zero- and few-shot LLMs on the same tasks. In this controlled setting, the lightweight models reach higher F1 with far fewer task-specific data. We treat this contrast as a comparison between learning regimes rather than a general verdict on LLMs. Our results show that careful input organisation supports sample-efficient learning of linguistic rules and reveals distinct learning signatures for trained lightweight models and prompted LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。