GPT-4可处理图文输入,表现接近人类水平。
GPT-4 Technical Report

- 基于Transformer的自回归预训练,支持多模态输入
- 在模拟律师考试中达前10%人类水平,事实性与行为对齐提升
- 通过可扩展基础设施实现性能精准预测,仅用千分之一算力
我们报告了GPT-4的开发,这是一个大规模多模态模型,可接受图像和文本输入并生成文本输出。尽管在许多现实场景中仍不及人类,但GPT-4在多个专业和学术基准测试中表现出人类水平的能力,例如在模拟律师考试中得分接近前10%的应试者。GPT-4是一个基于Transformer的模型,通过预训练以预测文档中的下一个标记。经过后训练对齐过程后,其在事实准确性及期望行为遵循性方面表现更优。本项目的核心是开发可在广泛规模下稳定运行的基础设施与优化方法,使我们能够基于最多仅使用GPT-4千分之一算力的模型,准确预测部分性能指标。
原文摘要 · Abstract (English)
We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs. While less capable than humans in many real-world scenarios, GPT-4 exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10% of test takers. GPT-4 is a Transformer-based model pre-trained to predict the next token in a document. The post-training alignment process results in improved performance on measures of factuality and adherence to desired behavior. A core component of this project was developing infrastructure and optimization methods that behave predictably across a wide range of scales. This allowed us to accurately predict some aspects of GPT-4's performance based on models trained with no more than 1/1,000th the compute of GPT-4.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。