arXiv:2508.19259cs.HCcs.CL2025-08被引 10

GPT-5在医疗、教学等领域表现优于GPT-4,展现更强领域适应能力。

Capabilities of GPT-5 across critical domains: Is it the next breakthrough?

  • 采用分模型架构,按任务需求优化性能
  • 专家评估显示GPT-5在四项任务中显著领先
  • 适合教育、医疗和学术研究场景使用

大语言模型的快速演进引发了对其在关键应用领域表现的探讨。GPT-4在推理、多模态和任务泛化方面取得进展,已应用于教育、临床诊断和学术写作,但存在缺陷。2025年8月发布的GPT-5采用系统化分模型架构,旨在实现任务特定优化。本研究是最早对GPT-4与GPT-5进行系统比较的实证分析之一,由20名语言学与临床领域的专家,基于预设标准,从五个维度(课程设计、作业评估、临床诊断、研究生成、伦理推理)评价模型输出。混合效应模型分析表明,GPT-5在课程设计、临床诊断、研究生成和伦理推理中显著优于GPT-4,而在作业评估上两者表现相当。结果凸显GPT-5作为情境敏感、领域专用工具的潜力,在教育、临床实践和学术研究中具有实际价值,并推动了伦理推理能力的发展。该研究为评估GPT-5演化能力与实用前景提供了早期实证依据。

原文摘要 · Abstract (English)

The accelerated evolution of large language models has raised questions about their comparative performance across domains of practical importance. GPT-4 by OpenAI introduced advances in reasoning, multimodality, and task generalization, establishing itself as a valuable tool in education, clinical diagnosis, and academic writing, though it was accompanied by several flaws. Released in August 2025, GPT-5 incorporates a system-of-models architecture designed for task-specific optimization and, based on both anecdotal accounts and emerging evidence from the literature, demonstrates stronger performance than its predecessor in medical contexts. This study provides one of the first systematic comparisons of GPT-4 and GPT-5 using human raters from linguistics and clinical fields. Twenty experts evaluated model-generated outputs across five domains: lesson planning, assignment evaluation, clinical diagnosis, research generation, and ethical reasoning, based on predefined criteria. Mixed-effects models revealed that GPT-5 significantly outperformed GPT-4 in lesson planning, clinical diagnosis, research generation, and ethical reasoning, while both models performed comparably in assignment assessment. The findings highlight the potential of GPT-5 to serve as a context-sensitive and domain-specialized tool, offering tangible benefits for education, clinical practice, and academic research, while also advancing ethical reasoning. These results contribute to one of the earliest empirical evaluations of the evolving capabilities and practical promise of GPT-5.

大模型评测GPT-5医疗AI教育应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。