化学文本嵌入基准测试,评估34个模型在化学领域的表现。
ChemTEB: Chemical Text Embedding Benchmark, an Overview of Embedding Models Performance & Efficiency on a Specific Domain
- 构建化学领域专用嵌入模型评测基准ChemTEB
- 34个模型在化学文献数据上性能对比分析
- 适合化学NLP研究者和模型开发者参考
近年来,语言模型的进展开启了信息检索与内容生成的新纪元,嵌入模型在提升数据表征效率与性能方面发挥关键作用。尽管如大规模文本嵌入基准(MTEB)等已标准化通用领域嵌入模型的评估,但在化学等专业领域仍存在空白,因该领域面临特有的语言与语义挑战。本文提出化学文本嵌入基准(ChemTEB),专为化学科学设计,针对化学文献与数据的特殊性,提供一套全面的任务集。通过在该基准上对34个开源与专有模型进行评估,揭示了当前方法在处理与理解化学信息方面的优劣。本工作旨在为研究社区提供一个标准化、领域特定的评估框架,推动更精确高效的化学相关NLP模型发展。同时,也揭示通用模型在特定领域中的表现。ChemTEB开源代码与数据,增强其可及性与实用性。
原文摘要 · Abstract (English)
Recent advancements in language models have started a new era of superior information retrieval and content generation, with embedding models playing an important role in optimizing data representation efficiency and performance. While benchmarks like the Massive Text Embedding Benchmark (MTEB) have standardized the evaluation of general domain embedding models, a gap remains in specialized fields such as chemistry, which require tailored approaches due to domain-specific challenges. This paper introduces a novel benchmark, the Chemical Text Embedding Benchmark (ChemTEB), designed specifically for the chemical sciences. ChemTEB addresses the unique linguistic and semantic complexities of chemical literature and data, offering a comprehensive suite of tasks on chemical domain data. Through the evaluation of 34 open-source and proprietary models using this benchmark, we illuminate the strengths and weaknesses of current methodologies in processing and understanding chemical information. Our work aims to equip the research community with a standardized, domain-specific evaluation framework, promoting the development of more precise and efficient NLP models for chemistry-related applications. Furthermore, it provides insights into the performance of generic models in a domain-specific context. ChemTEB comes with open-source code and data, contributing further to its accessibility and utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。