用五岁意大利儿童语言数据训练小型模型,验证语义理解需超越单纯数据量。
BAMBI: Developing Baby Language Models for Italian
- 基于儿童语言数据构建小规模意大利语语言模型
- 小模型语法能力接近大模型,但语义理解明显不足
- 适合研究语言习得机制与高效训练策略的学者
本文提出BAMBI(BAby language Models Boostrapped for Italian),一系列基于五岁意大利语儿童语言输入数据训练的婴儿语言模型(BabyLMs)。通过专门设计的评估基准测试模型,该基准考虑了模型接收的训练数据量。BAMBI模型与大型语言模型(LLM)及多模态语言模型(VLM)进行对比,以研究非语言信息对语言习得的贡献。评估结果与英语模型研究一致:有限训练数据虽能支持相对稳健的句法能力,但不足以促进语义理解。然而,尽管BAMBI模型在训练数据和计算资源上远少于LLM,其性能差距并未完全体现——大模型表现仅略优于小模型。这表明,除扩大训练资源外,数据筛选、多模态输入融入及课程学习等策略可能在塑造模型表现中起关键作用。
原文摘要 · Abstract (English)
This paper presents BAMBI (BAby language Models Boostrapped for Italian), a series of Baby Language Models (BabyLMs) trained on data that mimic the linguistic input received by a five-year-old Italian-speaking child. The BAMBI models are tested using a benchmark specifically designed to evaluate language models, which takes into account the amount of training input the models received. The BAMBI models are compared against a large language model (LLM) and a multimodal language model (VLM) to study the contribution of extralinguistic information for language acquisition. The results of our evaluation align with the existing literature on English language models, confirming that while reduced training data support the development of relatively robust syntactic competence, they are insufficient for fostering semantic understanding. However, the gap between the training resources (data and computation) of the BAMBI models and the LLMs is not fully reflected in their performance: despite LLMs' massive training, their performance is not much better than that of BAMBI models. This suggests that strategies beyond scaling training resources, such as data curation, inclusion of multimodal input, and other training strategies such as curriculum learning, could play a crucial role in shaping model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。