用大模型生成数据提升小领域模型性能,无需依赖专业术语库。
Enhancing Domain-Specific Encoder Models with LLM-Generated Data: How to Leverage Ontologies, and How to Do Without Them
- 用大模型生成数据扩充领域术语库,指导编码器预训练。
- 仅需少量论文摘要即可实现与大规模预训练相当的效果。
- 全自动流程适合低资源领域,无需人工构建知识图谱。
我们研究了在训练数据稀缺的特定领域(以入侵生物学为例)中,利用大语言模型生成数据对编码器模型进行持续预训练的有效性。通过将领域术语库与大模型生成数据结合,训练出基于术语库的嵌入模型以理解概念定义。为此,我们构建了一个专门用于评估入侵生物学模型性能的基准测试。结果表明,该方法显著优于标准大模型预训练。进一步探索了在缺乏完整术语库的领域中,通过从少量科学摘要中自动提取概念,并利用分布统计建立概念间关系的可行性。实验显示,该自动化方法仅需少量摘要即可达到与大规模掩码语言建模预训练相当的性能,实现了完全自动化的领域小模型增强流程,特别适用于低资源场景。
原文摘要 · Abstract (English)
We investigate the use of LLM-generated data for continual pretraining of encoder models in specialized domains with limited training data, using the scientific domain of invasion biology as a case study. To this end, we leverage domain-specific ontologies by enriching them with LLM-generated data and pretraining the encoder model as an ontology-informed embedding model for concept definitions. To evaluate the effectiveness of this method, we compile a benchmark specifically designed for assessing model performance in invasion biology. After demonstrating substantial improvements over standard LLM pretraining, we investigate the feasibility of applying the proposed approach to domains without comprehensive ontologies by substituting ontological concepts with concepts automatically extracted from a small corpus of scientific abstracts and establishing relationships between concepts through distributional statistics. Our results demonstrate that this automated approach achieves comparable performance using only a small set of scientific abstracts, resulting in a fully automated pipeline for enhancing domain-specific understanding of small encoder models that is especially suited for application in low-resource settings and achieves performance comparable to masked language modeling pretraining on much larger datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。