对比GPT-4o与DeepSeek R1在科学文本分类中的表现
Comparative Analysis of OpenAI GPT-4o and DeepSeek R1 for Scientific Text Categorization Using Prompt Engineering
- 设计专用提示工程评估方法,对比两模型分类能力
- 构建跨领域清洗科学论文数据集用于测试
- 首次系统评估DeepSeek R1在科学文本分类中的性能
本研究探讨大型语言模型如何通过提示工程对科学论文中的句子进行分类。我们采用两个先进的基于网络的模型——OpenAI的GPT-4o与DeepSeek R1,将句子归入预定义的关系类别。尽管DeepSeek R1已在技术报告中于基准数据集上测试过,但其在科学文本分类任务中的表现尚未被探索。为填补这一空白,我们提出一种专为此任务设计的评估方法,并整理了一个涵盖多个领域的清洗后科学论文数据集。该数据集为两模型的分类效果与一致性分析提供了平台。实验结果揭示了各模型在不同任务场景下的优劣。
原文摘要 · Abstract (English)
This study examines how large language models categorize sentences from scientific papers using prompt engineering. We use two advanced web-based models, GPT-4o (by OpenAI) and DeepSeek R1, to classify sentences into predefined relationship categories. DeepSeek R1 has been tested on benchmark datasets in its technical report. However, its performance in scientific text categorization remains unexplored. To address this gap, we introduce a new evaluation method designed specifically for this task. We also compile a dataset of cleaned scientific papers from diverse domains. This dataset provides a platform for comparing the two models. Using this dataset, we analyze their effectiveness and consistency in categorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。