arXiv:2409.06013cs.CLcs.CV2024-09中稿 · SpeD 2025

无需转录文本,用少量样本自动挖掘正负样本实现低资源语言关键词定位。

Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings

  • 引入少样本学习自动挖掘语音-图像匹配对,无需依赖转录文本。
  • 在英语上性能仅小幅下降,而在约鲁巴语上因挖掘不准导致性能显著降低。
  • 首次在真实低资源语言约鲁巴语上验证视觉提示关键词定位可行性。

给定一张图像查询,视觉提示关键词定位(VPKL)旨在从语音集合中找到与图像内容对应的词汇出现位置。该任务在缺乏转录文本的低资源语言(如未记录语言)中尤为有用。先前研究证明,可通过在配对图像和无标注语音上训练的视觉语音模型实现此任务,但所有实验均基于英语,且使用转录文本构建对比损失所需的正负样本对。本文提出一种少样本学习方案,可无需转录文本自动挖掘样本对。在英语上,性能仅小幅下降;首次在真实低资源语言约鲁巴语上进行测试,结果虽合理,但由于挖掘准确性较低,性能下降更明显。

原文摘要 · Abstract (English)

Given an image query, visually prompted keyword localisation (VPKL) aims to find occurrences of the depicted word in a speech collection. This can be useful when transcriptions are not available for a low-resource language (e.g. if it is unwritten). Previous work showed that VPKL can be performed with a visually grounded speech model trained on paired images and unlabelled speech. But all experiments were done on English. Moreover, transcriptions were used to get positive and negative pairs for the contrastive loss. This paper introduces a few-shot learning scheme to mine pairs automatically without transcriptions. On English, this results in only a small drop in performance. We also - for the first time - consider VPKL on a real low-resource language, Yoruba. While scores are reasonable, here we see a bigger drop in performance compared to using ground truth pairs because the mining is less accurate in Yoruba.

关键词定位低资源语言少样本学习视觉语音对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。