仅用标题摘要文本,就能识别论文质量高低。
Research Paper Quality Recognition Through Textual Feature Analysis

- 用文本嵌入+分类器分析标题摘要,区分高引用与撤稿论文。
- FastText+SVM达91.12%准确率,优于其他组合。
- 可视化与可解释性分析,适合学术诚信研究者参考。
科学知识与创新依赖于研究质量与可信度。然而,区分有影响力的高质量论文与存在缺陷的研究仍具挑战。本文提出一个基准,仅基于论文标题和摘要的文本特征,将研究论文分为两类:优质(高被引)与非优质(撤稿)。评估了SBERT、Word2Vec、FastText、USE和TF-IDF等多种嵌入技术,结合支持向量机(SVM)、随机森林和神经网络等分类器。贡献包括:(1)超参数透明性,(2)使用t-SNE的特征空间可视化,(3)基于SHAP的模型可解释性分析,(4)错误案例的深入分析。实验结果表明,采用SBERT嵌入的神经网络达到87.22%准确率,而FastText与SVM结合达到91.12%。这些发现凸显了文本信息在评估研究质量中的价值,并引发部署时的伦理思考。本研究为发展学术诚信工具、促进可信学术研究提供支持。
原文摘要 · Abstract (English)
Knowledge and innovations are shaped by using the quality and credibility of the scientific research. Yet, distinguishing between impactful, high-quality work and flawed studies remains a challenge. This paper introduces a benchmark for classifying research papers into two categories: good (highly cited) and non-good (retracted), using only textual features from titles and abstracts. We evaluate multiple embedding techniques, including SBERT, Word2Vec, FastText, USE, and TF-IDF, combined with classifiers such as Support Vector Machines (SVM), Random Forests, and Neural Networks. Our contributions include: (1) hyperparameter transparency, (2) feature space visualizations using t-SNE, (3) model interpretability analysis with SHAP, and (4) detailed examination of error cases. Experimental results show that a neural network with SBERT embeddings achieves 87.22\% accuracy, while FastText combined with SVM reaches 91.12\%. These findings highlight the value of textual information in assessing research quality, with ethical considerations for deployment. This work contributes toward the development of academic integrity tools that promote trustworthy scholarship.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。