arXiv:2409.02865cs.CLcs.CV2024-09被引 1

用图像引导语音模型识别关键词,助力低资源语言与语言习得研究

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

  • 通过无标签语音与图像配对训练视觉语音模型
  • 在尤鲁巴语等低资源语言中实现少样本高效学习
  • 揭示了多语言对认知偏见的影响差异,适合语言学与认知科学读者

本论文研究从无标注语音与图像配对中学习的视觉语音(VGS)模型,聚焦于低资源语言应用与人类语言习得理解。提出一种名为视觉提示关键词定位的新任务,利用图像检测并定位语音中的关键词。实验证明,该模型在尤鲁巴语等低资源语言的少样本学习场景中表现有效。此外,研究发现单语VGS模型存在互斥性偏见,而多语言训练并未像儿童那样显著影响此偏见,表明其认知机制不同于人类。

原文摘要 · Abstract (English)

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech paired with images. It focuses on applications for low-resource languages and understanding human language acquisition. We introduce a task called visually prompted keyword localisation to detect and localise keywords in speech using images. We demonstrate the effectiveness of VGS models in few-shot learning scenarios for low-resource languages like Yoruba. Additionally, we examine the mutual exclusivity bias in VGS models. Our monolingual VGS model exhibits this bias, but we found that multilingualism does not affect the bias in this VGS model similarly to what is observed in children.

视觉语音低资源语言认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。