用大模型重看科学概念演变史,揭示其继承与突破。
Computational conceptual history of scientific concepts: From early digital methods to LLMs
- 将大模型置于科学概念分析的数字方法长链中审视
- 发现大模型在语义变化检测上提升但仍受数据与训练影响
- 适合对科技史、计算人文感兴趣的学者参考
本文将大型语言模型(LLMs)置于科学史、哲学与社会学(HPSS)中计算概念分析的长期发展脉络中。首先回顾了大模型出现前的三类方法:早期数字方法、数字史中的分布性方法及词义变迁检测。梳理了语料构建、操作化与建模选择、评估与解释等方面的挑战与机遇。随后进入大模型时代,介绍其基础并回顾基于大模型的词义变迁检测与相关案例研究。重新审视此前的方法论问题,揭示语料构建、模型选择与训练数据、操作化权衡以及评估与解释在大模型工作流中的具体表现。
原文摘要 · Abstract (English)
This article situates large language models (LLMs) within the longer history of computational approaches to concept analysis in the history, philosophy, and sociology of science (HPSS). We examine what LLMs add to existing methods, how they inherit longstanding problems, and review recent case studies that employ them. In the first part, we reconstruct computational conceptual history before LLMs by bringing together three strands of work: early digital methods in HPSS, distributional approaches from digital history and related research, and lexical semantic change detection. We provide an overview of the main challenges and opportunities, focusing on corpus construction, operationalization and modelling choices, and evaluation and interpretation. In the second part, we turn to the era of LLMs, starting with a short introduction to LLMs before reviewing LLM-based work on lexical semantic change detection and relevant case studies in HPSS. We then revisit the earlier methodological questions, showing how issues of corpus construction, model choice and training data, operationalization trade-offs, and evaluation and interpretation play out in LLM-based workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。