arXiv:2606.16807cs.CL2026-06中稿 · EUSIPCO 2026 - 5 p…被引 1

用图像和语音无监督构建语音与文字的对应关系。

Connecting Speech to Words through Images

论文配图:Connecting Speech to Words through Images
图 1 · 摘自论文原文
  • 通过图像描述生成文字词汇,再匹配对应语音片段。
  • 在语音检索任务中超越强基线模型表现。
  • 适合低资源语言语音识别研究者参考。

在缺乏显式文本监督的情况下,如何学习书面词与发音之间的映射?我们提出一种基于视觉的方法,仅使用图像及其语音描述构建口语词汇表。首先,利用图像字幕系统提取图像中显著视觉概念对应的书面词;然后,对每个词,找出其图像字幕中包含该词的语音语段;最后,采用无监督词发现技术对齐这些语段,定位目标词的发音实例。最终获得与书面词关联的语音片段,整个过程无需任何文本标注。在语音词检索和关键词检测实验中,该方法优于强基线神经模型,且更具可解释性。结果证明了该方法在英语中的可行性,并为无转录文本的低资源语言研究提供了新思路。

原文摘要 · Abstract (English)

How can we learn the mapping between written words and their spoken counterparts in the absence of explicit textual supervision? We present a visually grounded method for building a vocabulary of spoken words using only images and their spoken descriptions. First, image captioning systems are used to build a vocabulary of written words representing salient visual concepts in the images. For each word, we then find utterances whose image captions contain that word. Then we use an unsupervised word discovery technique to align these utterances to locate instances of the target word. The result is spoken word segments that are linked to written words -- all accomplished without any text supervision. In spoken word retrieval and keyword spotting experiments, the proposed approach outperforms a strong neural baseline while being more interpretable. These results demonstrate the feasibility of the approach in English and motivate future work on low-resource languages without transcripts.

语音对齐无监督学习视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。