Gemini 2.5系列模型在推理、多模态和长上下文上全面突破,支持3小时视频处理。
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

- 融合长上下文、多模态与强推理能力,构建新型智能体工作流。
- Gemini 2.5 Pro在编码与推理基准上达到当前最佳性能。
- 提供从高算力到低延迟的全成本-性能覆盖,适合复杂任务自动化。
本文介绍Gemini 2.X系列模型:Gemini 2.5 Pro、Gemini 2.5 Flash,以及早期的Gemini 2.0 Flash和Flash-Lite。Gemini 2.5 Pro是迄今最强大的模型,在前沿编码与推理基准上达到当前最佳(SoTA)表现。除了卓越的编码与推理能力,它还具备思考型模型特性,擅长多模态理解,可处理长达3小时的视频内容。其长上下文、多模态与推理能力的结合,可解锁新型智能体工作流。Gemini 2.5 Flash在极低计算与延迟开销下实现优异推理能力;Gemini 2.0 Flash与Flash-Lite则在低延迟与低成本下保持高性能。整体上,Gemini 2.X系列覆盖了模型能力与成本之间的完整帕累托前沿,使用户能够探索复杂智能体问题求解的边界。
原文摘要 · Abstract (English)
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal understanding and it is now able to process up to 3 hours of video content. Its unique combination of long context, multimodal and reasoning capabilities can be combined to unlock new agentic workflows. Gemini 2.5 Flash provides excellent reasoning abilities at a fraction of the compute and latency requirements and Gemini 2.0 Flash and Flash-Lite provide high performance at low latency and cost. Taken together, the Gemini 2.X model generation spans the full Pareto frontier of model capability vs cost, allowing users to explore the boundaries of what is possible with complex agentic problem solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。