填补日语通用文本嵌入模型空白,用大模型生成数据训练新模型
Ruri: Japanese General Text Embeddings
- 用大模型合成数据训练日语嵌入模型
- 通过重排序器筛选数据并实现知识蒸馏
- 为日语NLP提供可落地的通用嵌入方案
我们报告了Ruri系列日语通用文本嵌入模型的开发。尽管近年来英语和多语言环境下的通用文本嵌入模型发展活跃,但日语领域的模型开发仍显不足,主要原因在于缺乏数据集和必要专业知识。本报告详细阐述了Ruri的开发过程,包括使用大语言模型生成的合成数据训练嵌入模型、构建用于数据过滤和知识蒸馏的重排序器,以及对最终通用文本嵌入模型的性能评估。
原文摘要 · Abstract (English)
We report the development of Ruri, a series of Japanese general text embedding models. While the development of general-purpose text embedding models in English and multilingual contexts has been active in recent years, model development in Japanese remains insufficient. The primary reasons for this are the lack of datasets and the absence of necessary expertise. In this report, we provide a detailed account of the development process of Ruri. Specifically, we discuss the training of embedding models using synthesized datasets generated by LLMs, the construction of the reranker for dataset filtering and knowledge distillation, and the performance evaluation of the resulting general-purpose text embedding models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。