arXiv:2510.06244cs.CLcs.AI2025-10

评测科学领域词向量与分词方法,寻找高效精准的文本表示方案。

Evaluating Embedding Frameworks for Scientific Domain

  • 构建科学领域专用的词向量与分词评估体系。
  • 在多个下游任务上测试不同算法表现,验证最优组合。
  • 适合从事科学文本NLP研究或模型优化的读者参考。

在特定领域数据中,选择最优的词表示算法尤为关键,因为同一词汇在不同领域和上下文中可能具有不同含义。尽管生成式AI与Transformer架构能生成上下文相关的嵌入表示,但其训练成本高,尤其从零开始预训练时耗时耗力。本文聚焦科学领域,旨在寻找适用于该领域的最优词表示算法与分词方法,并实现两个目标:1)确定可用于下游科学文本NLP任务的最佳表示与分词策略;2)建立一个全面的评估套件,以持续评估各类词表示与分词算法(包括未来新方法)。为此,我们设计了涵盖多个下游任务及对应数据集的评估框架,并利用该套件对多种算法进行测试,为科学领域自然语言处理提供可复用的基准工具。

原文摘要 · Abstract (English)

Finding an optimal word representation algorithm is particularly important in terms of domain specific data, as the same word can have different meanings and hence, different representations depending on the domain and context. While Generative AI and transformer architecture does a great job at generating contextualized embeddings for any given work, they are quite time and compute extensive, especially if we were to pre-train such a model from scratch. In this work, we focus on the scientific domain and finding the optimal word representation algorithm along with the tokenization method that could be used to represent words in the scientific domain. The goal of this research is two fold: 1) finding the optimal word representation and tokenization methods that can be used in downstream scientific domain NLP tasks, and 2) building a comprehensive evaluation suite that could be used to evaluate various word representation and tokenization algorithms (even as new ones are introduced) in the scientific domain. To this end, we build an evaluation suite consisting of several downstream tasks and relevant datasets for each task. Furthermore, we use the constructed evaluation suite to test various word representation and tokenization algorithms.

词向量科学文本评估套件分词方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。