GPT-4o支持多模态输入输出,响应快、成本低,且在非英语文本和音视频理解上表现更优。
GPT-4o System Card
- 统一神经网络处理文本、图像、音频与视频输入输出
- 语音响应最快232毫秒,平均320毫秒,接近人类对话速度
- 非英语文本性能显著提升,API成本降低50%,适合多模态应用
GPT-4o 是一个自回归的全模态模型,可接受文本、音频、图像和视频任意组合输入,并生成文本、音频和图像任意组合输出。它在文本、视觉和音频上端到端训练,所有输入输出由同一神经网络处理。GPT-4o 语音输入响应最快仅需232毫秒,平均320毫秒,接近人类对话反应时间。其英文文本与代码能力媲美GPT-4 Turbo,非英语文本性能显著提升,同时推理速度更快,API成本降低50%。在视觉与音频理解方面优于现有模型。为保障安全,我们公开了包含准备度框架评估的系统卡片,涵盖语音到语音、文本与图像能力的安全评估,以及第三方对危险能力的测试,讨论了其可能的社会影响。
原文摘要 · Abstract (English)
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's trained end-to-end across text, vision, and audio, meaning all inputs and outputs are processed by the same neural network. GPT-4o can respond to audio inputs in as little as 232 milliseconds, with an average of 320 milliseconds, which is similar to human response time in conversation. It matches GPT-4 Turbo performance on text in English and code, with significant improvement on text in non-English languages, while also being much faster and 50\% cheaper in the API. GPT-4o is especially better at vision and audio understanding compared to existing models. In line with our commitment to building AI safely and consistent with our voluntary commitments to the White House, we are sharing the GPT-4o System Card, which includes our Preparedness Framework evaluations. In this System Card, we provide a detailed look at GPT-4o's capabilities, limitations, and safety evaluations across multiple categories, focusing on speech-to-speech while also evaluating text and image capabilities, and measures we've implemented to ensure the model is safe and aligned. We also include third-party assessments on dangerous capabilities, as well as discussion of potential societal impacts of GPT-4o's text and vision capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。