arXiv:2512.08057cs.AI2025-12综述被引 1

对比ChatGPT与DeepSeek在教育科研中的表现与体验。

Large Language Models for Education and Research: An Empirical and User Survey-based Analysis

  • 通过实验与用户调研,评估两模型在教学科研中的表现
  • ChatGPT文本理解更强,DeepSeek编程效率更优
  • 两者均能准确诊断疾病并解复杂数学题,适合教育研究者使用

预训练大语言模型(LLMs)在多个领域取得显著进展,教育与科研成为其最具影响力的应用方向。当前先进模型如ChatGPT与DeepSeek在数学、科学、医学、文学及编程方面表现突出。本研究通过技术背景分析、实证实验与真实用户调研,全面评估这两款模型在教育与科研场景下的性能。评估涵盖模型准确性、计算效率与用户体验之间的权衡。实验对比了模型在文本生成、编程与专业问题求解方面的表现。结果表明,ChatGPT在通用语言理解与文本生成上表现更佳,而DeepSeek因面向效率设计,在编程任务中优势明显。此外,两款模型均能提供医学诊断的准确输出,并有效解决复杂数学问题。结合定量结果,对师生与研究人员的问卷调查揭示了模型的实际价值与局限,深化了对大模型在教育科研中作用的理解。

原文摘要 · Abstract (English)

Pretrained Large Language Models (LLMs) have achieved remarkable success across diverse domains, with education and research emerging as particularly impactful areas. Among current state-of-the-art LLMs, ChatGPT and DeepSeek exhibit strong capabilities in mathematics, science, medicine, literature, and programming. In this study, we present a comprehensive evaluation of these two LLMs through background technology analysis, empirical experiments, and a real-world user survey. The evaluation explores trade-offs among model accuracy, computational efficiency, and user experience in educational and research affairs. We benchmarked these LLMs performance in text generation, programming, and specialized problem-solving. Experimental results show that ChatGPT excels in general language understanding and text generation, while DeepSeek demonstrates superior performance in programming tasks due to its efficiency-focused design. Moreover, both models deliver medically accurate diagnostic outputs and effectively solve complex mathematical problems. Complementing these quantitative findings, a survey of students, educators, and researchers highlights the practical benefits and limitations of these models, offering deeper insights into their role in advancing education and research.

大模型教育应用用户调研科研助手

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。