首个专为人文社科设计的因果语言模型,性能接近通用模型但训练量少100倍。
SHARE: Social-Humanities AI for Research and Education
- 基于人文社科文本全量预训练,构建首个因果语言模型家族
- 在自研评测中性能逼近通用模型,仅用其1%训练数据
- 界面不生成内容,保障学术批判性与人文研究规范
本技术报告介绍SHARE系列基础模型及MIRROR用户界面。SHARE模型是首个完全由社会科学与人文学科(SSH)数据预训练并专为该领域设计的因果语言模型。在自研的SSH Cloze评测中,其表现接近使用100倍训练数据的通用模型(Phi-4)。MIRROR界面旨在支持人文社科文本的审阅,同时保持批判性参与。通过原型设计一个不生成文本的生成式AI界面,提出一种在不损害人文研究原则的前提下利用SHARE模型能力的新方式。
原文摘要 · Abstract (English)
This intermediate technical report introduces the SHARE family of base models and the MIRROR user interface. The SHARE models are the first causal language models fully pretrained by and for the social sciences and humanities (SSH). Their performance in modelling SSH texts is close to that of general purpose models (Phi-4) which use 100 times more tokens, as shown by our custom SSH Cloze benchmark. The MIRROR user interface is designed for reviewing text inputs from the SSH disciplines while preserving critical engagement. By prototyping a generative AI interface that does not generate any text, we propose a way to harness the capabilities of the SHARE models without compromising the integrity of SSH principles and norms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。