arXiv:2410.01188cs.CL2024-10EMNLP被引 17

通过智能筛选词汇子集,提升特定领域大模型表现

Gold Panning in Vocabulary: An Adaptive Method for Vocabulary Expansion of Domain-Specific LLMs

  • 基于领域词汇自动识别高价值词,动态扩展词库
  • 在三个中文数据集上验证,任务性能显著提升
  • 适合需要高效适配专业领域的模型开发者

尽管大语言模型具备出色的生成能力,但在专业领域常因缺乏特定知识而表现不佳。现有研究通常在微调前扩展词汇表以减少序列长度、提高解码效率,却未深入分析不同领域下词汇扩展的实际效果。我们的初步研究发现,仅使用部分词汇即可取得更优性能。受此启发,本文提出 VEGAD——一种自适应方法,可从给定领域词汇中自动筛选出有价值词语。实验在三个中文数据集上验证了该方法的有效性。综合分析表明,选择最优词汇子集进行扩展,不仅提升了领域任务表现,也改善了通用任务性能,展现了 VEGAD 的潜力。

原文摘要 · Abstract (English)

While Large Language Models (LLMs) demonstrate impressive generation abilities, they frequently struggle when it comes to specialized domains due to their limited domain-specific knowledge. Studies on domain-specific LLMs resort to expanding the vocabulary before fine-tuning on domain-specific corpus, aiming to decrease the sequence length and enhance efficiency during decoding, without thoroughly investigating the results of vocabulary expansion to LLMs over different domains. Our pilot study reveals that expansion with only a subset of the entire vocabulary may lead to superior performance. Guided by the discovery, this paper explores how to identify a vocabulary subset to achieve the optimal results. We introduce VEGAD, an adaptive method that automatically identifies valuable words from a given domain vocabulary. Our method has been validated through experiments on three Chinese datasets, demonstrating its effectiveness. Additionally, we have undertaken comprehensive analyses of the method. The selection of a optimal subset for expansion has shown to enhance performance on both domain-specific tasks and general tasks, showcasing the potential of VEGAD.

词汇扩展大模型领域适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。