对比三大模型在科学计算中的表现,发现推理优化模型更优。
DeepSeek vs. ChatGPT vs. Claude: A Comparative Study for Scientific Computing and Scientific Machine Learning Tasks
- 用推理增强方法评估模型解决科学计算难题的能力
- 推理模型在复杂问题上显著优于普通模型,ChatGPT o3-mini-high速度最快
- 适合关注科学计算与机器学习融合的科研人员参考
大型语言模型(LLMs)已成为解决各类问题的强大工具,尤其在科学计算领域,如求解偏微分方程(PDEs)。然而,不同模型展现出各异的优势与偏好,导致性能差异。本文对比了最先进的LLMs——DeepSeek、ChatGPT和Claude及其推理优化版本在应对计算挑战时的表现。具体评估其在传统科学计算数值问题上的能力,以及利用科学机器学习技术求解基于PDE的问题。所有实验设计均要求非平凡决策,例如为神经算子学习定义合适的输入函数空间。结果表明,推理及混合推理模型在解决难题时始终显著优于非推理模型,其中ChatGPT o3-mini-high普遍具备最快推理速度。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have emerged as powerful tools for tackling a wide range of problems, including those in scientific computing, particularly in solving partial differential equations (PDEs). However, different models exhibit distinct strengths and preferences, resulting in varying levels of performance. In this paper, we compare the capabilities of the most advanced LLMs--DeepSeek, ChatGPT, and Claude--along with their reasoning-optimized versions in addressing computational challenges. Specifically, we evaluate their proficiency in solving traditional numerical problems in scientific computing as well as leveraging scientific machine learning techniques for PDE-based problems. We designed all our experiments so that a non-trivial decision is required, e.g. defining the proper space of input functions for neural operator learning. Our findings show that reasoning and hybrid-reasoning models consistently and significantly outperform non-reasoning ones in solving challenging problems, with ChatGPT o3-mini-high generally offering the fastest reasoning speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。