收集9.4万条真实用户使用大模型的案例,揭示其应用分布与人群特征。
REALM: A Dataset of Real-World LLM Use Cases
- 从Reddit和新闻中收集9.4万条真实用例,覆盖多样场景
- 发现不同职业群体偏好使用大模型的不同类型应用
- 为研究大模型社会影响提供可扩展的真实数据基础
大型语言模型(如GPT系列)已推动显著的产业应用,带来经济与社会变革。然而,对其真实世界应用的全面理解仍有限。为此,我们推出了REALM数据集,包含超过9.4万条来自Reddit和新闻文章的LLM使用案例。该数据集捕捉了两大关键维度:大模型的多样化应用场景及其使用者的人口统计特征。它对大模型应用进行分类,并探究用户职业与其使用场景之间的关联。通过整合真实世界数据,REALM揭示了大模型在不同领域的采纳情况,为未来研究其不断演变的社会角色提供了坚实基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs), such as the GPT series, have driven significant industrial applications, leading to economic and societal transformations. However, a comprehensive understanding of their real-world applications remains limited. To address this, we introduce REALM, a dataset of over 94,000 LLM use cases collected from Reddit and news articles. REALM captures two key dimensions: the diverse applications of LLMs and the demographics of their users. It categorizes LLM applications and explores how users' occupations relate to the types of applications they use. By integrating real-world data, REALM offers insights into LLM adoption across different domains, providing a foundation for future research on their evolving societal roles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。