arXiv:2601.10088cs.AI2026-01被引 22

分析100万亿令牌数据,揭示大模型真实使用场景与用户行为模式。

State of AI: An Empirical 100 Trillion Token Study with OpenRouter

  • 基于OpenRouter平台分析超100万亿令牌交互数据。
  • 发现开源模型广泛采用,创意角色扮演和编程辅助需求远超预期。
  • 识别出早期用户留存率极高的'灰姑娘效应',对产品设计有重要启示。

过去一年标志着大语言模型(LLMs)演化与实际应用的重要转折点。随着2024年12月5日首个广泛采用的推理模型o1发布,领域从单次模式生成转向多步推理性推理,加速了部署、实验及新型应用的发展。然而,这一转变迅速推进,而我们对模型实际使用情况的实证理解却滞后。本研究利用OpenRouter平台——一个覆盖多种LLM的AI推理服务——分析了超过100万亿令牌的真实世界LLM交互数据,涵盖任务类型、地理分布与时间维度。研究发现,开源权重模型被大量采用,创意角色扮演(远超预期的生产力任务)与编码辅助类别异常流行,且代理式推理趋势上升。此外,留存分析揭示了核心用户群:早期用户留存时间显著长于后续群体,我们将其称为“灰姑娘效应”。这些发现表明,开发者与终端用户在真实环境中的互动复杂多样。本文讨论其对模型构建者、AI开发者与基础设施提供者的启示,并提出应以数据驱动的方式优化大模型系统的设计与部署。

原文摘要 · Abstract (English)

The past year has marked a turning point in the evolution and real-world use of large language models (LLMs). With the release of the first widely adopted reasoning model, o1, on December 5th, 2024, the field shifted from single-pass pattern generation to multi-step deliberation inference, accelerating deployment, experimentation, and new classes of applications. As this shift unfolded at a rapid pace, our empirical understanding of how these models have actually been used in practice has lagged behind. In this work, we leverage the OpenRouter platform, which is an AI inference provider across a wide variety of LLMs, to analyze over 100 trillion tokens of real-world LLM interactions across tasks, geographies, and time. In our empirical study, we observe substantial adoption of open-weight models, the outsized popularity of creative roleplay (beyond just the productivity tasks many assume dominate) and coding assistance categories, plus the rise of agentic inference. Furthermore, our retention analysis identifies foundational cohorts: early users whose engagement persists far longer than later cohorts. We term this phenomenon the Cinderella "Glass Slipper" effect. These findings underscore that the way developers and end-users engage with LLMs "in the wild" is complex and multifaceted. We discuss implications for model builders, AI developers, and infrastructure providers, and outline how a data-driven understanding of usage can inform better design and deployment of LLM systems.

大模型应用用户行为数据洞察

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。