首个越南语信息检索基准,提升中文之外的多语言研究。
Advancing Vietnamese Information Retrieval with Learning Objective and Benchmark
- 设计新损失函数优化越南语嵌入模型
- 构建首个越南语信息检索基准数据集
- 适合多语言NLP与检索系统研究者
随着自然语言处理技术的快速发展,众多语言模型被用于多种任务。其中信息检索(IR)至关重要,需模型准确召回相关文档。尽管在检索增强生成(RAG)等实际应用中意义重大,当前仍缺乏越南语信息检索的基准评测体系,导致现有越南语嵌入模型难以评估与对比,阻碍了越南语自然语言处理研究进展。为此,本文提出一个专注于检索与重排序任务的越南语信息检索新基准,并设计基于InfoNCE损失的新型目标函数,以提升模型在信息检索任务中的表现。同时,我们分析了温度超参数在两种目标函数中的影响,为模型调优提供依据。
原文摘要 · Abstract (English)
With the rapid development of natural language processing, many language models have been invented for multiple tasks. One important task is information retrieval (IR), which requires models to retrieve relevant documents. Despite its importance in many real-life applications, especially in retrieval augmented generation (RAG) systems, this task lacks Vietnamese benchmarks. This situation causes difficulty in assessing and comparing many existing Vietnamese embedding language models on the task and slows down the advancement of Vietnamese natural language processing (NLP) research. In this work, we aim to provide the Vietnamese research community with a new benchmark for information retrieval, which mainly focuses on retrieval and reranking tasks. Furthermore, we also present a new objective function based on the InfoNCE loss function, which is used to train our Vietnamese embedding model. Our function aims to be better than the origin in information retrieval tasks. Finally, we analyze the effect of temperature, a hyper-parameter in both objective functions, on the performance of text embedding models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。