首次揭示大模型如何用特定特征编码科研质量,助力理解AI评价机制。
How Do LLMs Encode Scientific Quality? An Empirical Study Using Monosemantic Features from Sparse Autoencoders
- 用稀疏自编码器提取单一语义特征,分析大模型内部质量表征
- 四类特征可预测引用量、期刊影响因子等科研质量指标
- 适合研究AI评估机制、科研评价智能化的学者参考
近年来,生成式AI尤其是大语言模型(LLMs)在科学工作评估与生成中应用日益广泛。尽管已有研究显示LLMs可在一定程度上依据感知质量评估研究,但其内部实现机制仍不清晰。本文首次通过稀疏自编码器提取的单义特征,探究LLMs如何编码科研质量概念。我们在不同实验设置下提取特征,并在三个科研质量相关任务中评估其预测能力:预测论文引用量、期刊SJR和期刊h指数。结果表明,LLMs编码了多维度科研质量特征。特别识别出四类反复出现的特征:1)反映研究方法的特征;2)与发表类型相关的特征,其中综述类文章通常具有更高影响力;3)与高影响力研究领域及技术相关的特征;4)对应特定科学术语的特征。这些发现为理解大模型如何内化科研质量概念迈出重要一步。
原文摘要 · Abstract (English)
In recent years, there has been a growing use of generative AI, and large language models (LLMs) in particular, to support both the assessment and generation of scientific work. Although some studies have shown that LLMs can, to a certain extent, evaluate research according to perceived quality, our understanding of the internal mechanisms that enable this capability remains limited. This paper presents the first study that investigates how LLMs encode the concept of scientific quality through relevant monosemantic features extracted using sparse autoencoders. We derive such features under different experimental settings and assess their ability to serve as predictors across three tasks related to research quality: predicting citation count, journal SJR, and journal h-index. The results indicate that LLMs encode features associated with multiple dimensions of scientific quality. In particular, we identify four recurring types of features that capture key aspects of how research quality is represented: 1) features reflecting research methodologies; 2) features related to publication type, with literature reviews typically exhibiting higher impact; 3) features associated with high-impact research fields and technologies; and 4) features corresponding to specific scientific jargons. These findings represent an important step toward understanding how LLMs encapsulate concepts related to research quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。