首个跨平台对话数据集,揭示大模型真实使用差异。
ShareChat: A Dataset of Chatbot Conversations in the Wild
- 采集5大平台14万+对话,保留原文引用、推理痕迹等原始功能
- 发现不同平台在意图满足度、引用策略、响应延迟上存在显著差异
- 适合研究多平台交互行为、模型实际表现与系统设计影响
当前学术评测通过统一的纯文本接口评估大语言模型,掩盖了不同商业平台独特设计对真实用户行为与系统性能的影响。为弥合这一差距,我们提出ShareChat,首个大规模对话语料库,包含来自ChatGPT、Perplexity、Grok、Gemini和Claude的142,808次对话(共660,293轮),数据源自公开分享的链接,覆盖95种语言,时间跨度为2023年4月至2025年10月。该数据集保留了各平台原生功能,如引用、思维链与代码片段。为验证其评估价值,我们开展三项案例研究:对话完整性分析,比较平台间意图满足度;来源溯源分析,对比搜索增强型系统的引用策略;时间动态分析,揭示响应延迟的异构演化模式。这些研究问题在单一平台或去功能化语料中无法实现。数据集已公开。
原文摘要 · Abstract (English)
By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system performance. To bridge this gap, we present ShareChat, the first large-scale corpus of 142,808 conversations (660,293 turns) collected from publicly shared URLs on ChatGPT, Perplexity, Grok, Gemini, and Claude. ShareChat preserves native platform affordances, including citations, thinking traces, and code artifacts, across 95 languages and the period from April 2023 to October 2025, complementing existing corpora that homogenize these interactions. To demonstrate the dataset's evaluative utility, we present three case studies: a conversation completeness analysis assessing cross-platform differences in intent satisfaction, a source grounding analysis comparing citation strategies between search-augmented systems, and a temporal analysis revealing divergent response latency dynamics. Together, these analyses demonstrate research questions that are inaccessible to single-platform or stripped-affordance corpora. The dataset is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。