用天文高能论文训练BERT,让模型理解科学概念的深层含义
Astro-HEP-BERT: A bidirectional language model for studying the meanings of concepts in astrophysics and high energy physics
- 基于arXiv海量论文微调BERT,生成科学概念上下文嵌入
- 在小数据集上表现媲美从头训练的大模型,效果显著
- 适合研究科学史、哲学与社会学的学者快速获取语义分析工具
本文提出Astro-HEP-BERT,一种基于Transformer的双向语言模型,专为生成天体物理与高能物理领域概念的上下文词嵌入(CWEs)而设计。该模型以通用预训练BERT为基础,使用我自建的Astro-HEP Corpus进行三轮微调,该语料库包含从超过60万篇arXiv学术文章中提取的2184万段文本,均属于天体物理或高能物理领域。研究表明,将通用语言模型适配特定科学领域,在科学史、哲学与社会学(HPSS)研究中具有可行性和有效性。整个训练过程仅使用免费代码、预训练权重和公开文本,在一台配备M2芯片和96GB内存的MacBook Pro上完成。初步评估显示,Astro-HEP-BERT生成的CWEs在领域特定词义消歧、语义归纳及语义变迁分析等任务中,性能可媲美从头训练于更大数据集的专用BERT模型,表明对通用模型进行领域微调是高效且低成本的策略,适用于无须大规模训练即可实现高性能的HPSS研究。
原文摘要 · Abstract (English)
I present Astro-HEP-BERT, a transformer-based language model specifically designed for generating contextualized word embeddings (CWEs) to study the meanings of concepts in astrophysics and high-energy physics. Built on a general pretrained BERT model, Astro-HEP-BERT underwent further training over three epochs using the Astro-HEP Corpus, a dataset I curated from 21.84 million paragraphs extracted from more than 600,000 scholarly articles on arXiv, all belonging to at least one of these two scientific domains. The project demonstrates both the effectiveness and feasibility of adapting a bidirectional transformer for applications in the history, philosophy, and sociology of science (HPSS). The entire training process was conducted using freely available code, pretrained weights, and text inputs, completed on a single MacBook Pro Laptop (M2/96GB). Preliminary evaluations indicate that Astro-HEP-BERT's CWEs perform comparably to domain-adapted BERT models trained from scratch on larger datasets for domain-specific word sense disambiguation and induction and related semantic change analyses. This suggests that retraining general language models for specific scientific domains can be a cost-effective and efficient strategy for HPSS researchers, enabling high performance without the need for extensive training from scratch.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。