用语音查法语文本,不用昂贵翻译管道
Cross-lingual Matryoshka Representation Learning across Speech and Text
- 构建双语语音-文本嵌入模型,实现方言语音到法语文本的高效检索
- 在沃尔夫语语音查询下,法语文本召回率超80%,媲美传统翻译链
- 适合低资源语言跨模态检索,尤其关注语音转文本的效率优化
低资源语言使用者面临语言壁垒(多数知识为少数主流语言)和模态壁垒(信息多为文本,但许多语言以口语为主)。本文针对法语-沃尔夫语,首次训练了双语语音-文本马特里什卡嵌入模型,无需依赖昂贵的语音识别-翻译流水线,即可实现沃尔夫语语音查询对法语文本的高效检索。研究构建了大规模数据清洗流程与新基准,比较不同建模策略,发现冻结文本马特里什卡模型内的模态融合效果最佳。尽管仅用于检索任务训练,模型在语音意图识别等下游任务上表现良好,表明其学习到了通用语义表征。最后分析了不同马特里什卡层级与秩下的成本-精度权衡,发现信息主要集中于少数组件中,暗示存在显著的效率提升空间。
原文摘要 · Abstract (English)
Speakers of under-represented languages face both a language barrier, as most online knowledge is in a few dominant languages, and a modality barrier, since information is largely text-based while many languages are primarily oral. We address this for French-Wolof by training the first bilingual speech-text Matryoshka embedding model, enabling efficient retrieval of French text from Wolof speech queries without relying on a costly ASR-translation pipelines. We introduce large-scale data curation pipelines and new benchmarks, compare modeling strategies, and show that modality fusion within a frozen text Matryoshka model performs best. Although trained only for retrieval, the model generalizes well to other tasks, such as speech intent detection, indicating the learning of general semantic representations. Finally, we analyze cost-accuracy trade-offs across Matryoshka dimensions and ranks, showing that information is concentrated only in a few components, suggesting potential for efficiency improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。