社会科学研究者选大模型,应从小而开源的开始,并验证全流程可靠性。
Selecting Language Models for Social Science: Start Small, Start Open, and Validate
- 优先选择小规模、开源模型,降低使用门槛和透明度风险。
- 强调事后验证计算结果的可复制性,而非仅依赖预设基准测试。
- 建议构建限定范围的基准,确保整个分析流程可重现。
目前社会科学家可选择的大型预训练语言模型(LLMs)已达数千个。如何筛选?本文以有效性、可靠性、可重复性和可复制性为指导原则,探讨模型开源性、模型规模、训练数据及架构与微调方式的影响。尽管事前有效性测试(如基准评估)常被重视,但本文主张社会科学家无法完全避免对计算测量方法进行事后验证。特别是可复制性,是选择模型的关键依据——要可靠复现研究结论,就必须能稳定再现特定任务。为此,我们建议从较小、开源模型入手,并构建限定范围的基准,以验证整个计算流程的有效性。
原文摘要 · Abstract (English)
Currently, there are thousands of large pretrained language models (LLMs) available to social scientists. How do we select among them? Using validity, reliability, reproducibility, and replicability as guides, we explore the significance of: (1) model openness, (2) model footprint, (3) training data, and (4) model architectures and fine-tuning. While ex-ante tests of validity (i.e., benchmarks) are often privileged in these discussions, we argue that social scientists cannot altogether avoid validating computational measures (ex-post). Replicability, in particular, is a more pressing guide for selecting language models. Being able to reliably replicate a particular finding that entails the use of a language model necessitates reliably reproducing a task. To this end, we propose starting with smaller, open models, and constructing delimited benchmarks to demonstrate the validity of the entire computational pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。