arXiv:2412.03223cs.CL2024-12被引 22

通过精调数据与混合任务训练,提升文本检索精度。

Linq-Embed-Mistral Technical Report

  • 基于E5-Mistral和Mistral-7B改进,定制化数据构建与筛选方法。
  • MTEB平均分68.2,检索任务排名第一,得分60.2。
  • 支持4位量化加速评估,适合追求高效高精度检索的场景。

本报告研究了通过先进数据精炼技术提升文本检索性能的方法。我们基于E5-mistral和Mistral-7B-v0.1模型开发了Linq-Embed-Mistral,聚焦于针对不同任务高度定制的数据构造、数据过滤及负样本挖掘方法,应用于现有基准数据集和由大语言模型生成的定制化合成数据集。Linq-Embed-Mistral在MTEB基准测试中表现优异(截至2024年5月29日),在56个数据集上平均得分为68.2,检索任务排名位居榜首,得分为60.2。该性能凸显其在提升搜索精确性与可靠性方面的卓越能力。我们的贡献包括显著提升模型在基准与合成数据集上表现的数据精炼方法、同质任务排序与混合任务微调技术以增强泛化性与稳定性,以及采用4比特精度与轻量检索评估集的简化评估流程,可在不损失准确率的前提下加速验证。

原文摘要 · Abstract (English)

This report explores the enhancement of text retrieval performance using advanced data refinement techniques. We develop Linq-Embed-Mistral\footnote{\url{https://huggingface.co/Linq-AI-Research/Linq-Embed-Mistral}} by building on the E5-mistral and Mistral-7B-v0.1 models, focusing on sophisticated data crafting, data filtering, and negative mining methods, which are highly tailored to each task, applied to both existing benchmark dataset and highly tailored synthetic dataset generated via large language models (LLMs). Linq-Embed-Mistral excels in the MTEB benchmarks (as of May 29, 2024), achieving an average score of 68.2 across 56 datasets, and ranks 1st among all models for retrieval tasks on the MTEB leaderboard with a performance score of 60.2. This performance underscores its superior capability in enhancing search precision and reliability. Our contributions include advanced data refinement methods that significantly improve model performance on benchmark and synthetic datasets, techniques for homogeneous task ordering and mixed task fine-tuning to enhance model generalization and stability, and a streamlined evaluation process using 4-bit precision and a light retrieval evaluation set, which accelerates validation without sacrificing accuracy.

文本检索数据精炼模型优化4位量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。