无需转写,仅靠图像就能让英文词对应到印地语发音。
Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

- 用图像自动生成英文标签,通过自监督语音特征对齐
- 在关键词检测任务中超越注意力神经模型,准确率达83.6%
- 适合低资源语言数据构建,无需人工标注
在许多低资源场景中,收集语音数据极为困难。一种可行方法是让说话者描述图像。本文基于包含印地语口语描述的图像数据集,研究如何将英文书面词映射为印地语中的实际发音。以往工作采用端到端多模态神经模型,而本文提出一种基于自监督语音表征的简化对齐方法。利用现成图像字幕系统自动获取图像对应的英文标签,再通过自监督特征对齐同一关键词的印地语语句,聚合对齐证据以识别重复出现的语音片段。实验表明,在关键词定位与检测任务中,该方法优于先前的注意力神经模型。引入负样本进一步提升了性能。结果证明,跨语言词到语音的映射可直接从视觉对齐中学习,无需转录或显式训练。
原文摘要 · Abstract (English)
In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speech data? Given a dataset of images with Hindi spoken captions, we consider how we can map a written English keyword to spoken realisations of that word in Hindi. Previous work trained end-to-end multimodal neural models. Instead, we explore a simpler alignment-based approach built on self-supervised speech representations. Written English tags are automatically obtained from images using off-the-shelf image captioning systems. Hindi utterances associated with the same keyword are then aligned (using self-supervised features), and alignment evidence is aggregated to identify recurring speech segments corresponding to the target word. Experiments evaluating keyword spotting and localization show that our alignment-based approach outperforms a previous attention-based neural model. We also show the benefit of incorporating negative examples during alignment. Our work demonstrates that cross-lingual word-to-speech mappings can be learned directly from visual grounding without transcriptions or explicit model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。