arXiv:2411.14073cs.CLphysics.hist-ph2024-11被引 4

用词嵌入分析科学概念演变,揭示'普朗克'含义三十年变迁

Meaning at the Planck scale? Contextualized word embeddings for doing history, philosophy, and sociology of science

  • 基于领域预训练的BERT模型识别科学术语多义性
  • 在2900个标注样本中准确区分'普朗克'的多种含义
  • 适合研究科学话语历史演变的学者使用

本文探索上下文词嵌入(CWEs)作为历史、哲学与科学社会学(HPSS)研究新工具的潜力,用于分析科学概念意义的演变。以'普朗克'为案例,评估五种基于BERT的模型,包括在60万篇天体物理与高能物理论文(2184万段落)上训练的自研模型Astro-HEP-BERT。构建两个标注数据集:(1) 从1500段落中抽取的2900个'普朗克'实例组成的Astro-HEP-Planck Corpus;(2) 包含885段落、1186个标注实例的物理类维基数据集。结果表明,领域适配模型在消歧、意义预测和语义聚类方面优于通用模型,新提出的纯净度指标验证了其性能。该方法还揭示了未标注数据中'普朗克'含义在过去三十年的演变,突显普朗克空间任务成为主导意义。研究强调领域预训练的重要性,并展示适应预训练模型在HPSS研究中的成本效益。通过提供可扩展、可迁移的科学概念意义建模方法,为探究科学话语的社会历史动态开辟新路径。

原文摘要 · Abstract (English)

This paper explores the potential of contextualized word embeddings (CWEs) as a new tool in the history, philosophy, and sociology of science (HPSS) for studying contextual and evolving meanings of scientific concepts. Using the term "Planck" as a test case, I evaluate five BERT-based models with varying degrees of domain-specific pretraining, including my custom model Astro-HEP-BERT, trained on the Astro-HEP Corpus, a dataset containing 21.84 million paragraphs from 600,000 articles in astrophysics and high-energy physics. For this analysis, I compiled two labeled datasets: (1) the Astro-HEP-Planck Corpus, consisting of 2,900 labeled occurrences of "Planck" sampled from 1,500 paragraphs in the Astro-HEP Corpus, and (2) a physics-related Wikipedia dataset comprising 1,186 labeled occurrences of "Planck" across 885 paragraphs. Results demonstrate that the domain-adapted models outperform the general-purpose ones in disambiguating the target term, predicting its known meanings, and generating high-quality sense clusters, as measured by a novel purity indicator I developed. Additionally, this approach reveals semantic shifts in the target term over three decades in the unlabeled Astro-HEP Corpus, highlighting the emergence of the Planck space mission as a dominant sense. The study underscores the importance of domain-specific pretraining for analyzing scientific language and demonstrates the cost-effectiveness of adapting pretrained models for HPSS research. By offering a scalable and transferable method for modeling the meanings of scientific concepts, CWEs open up new avenues for investigating the socio-historical dynamics of scientific discourses.

词嵌入科学史语义演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。