arXiv:2605.29018cs.AIcs.CL2026-05被引 1

分析1.2万用户对话,发现人用大模型习惯难改变,活跃者更高效。

Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild

论文配图:Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild
图 1 · 摘自论文原文
  • 追踪1.2万用户长期对话,发现个体行为变化微弱
  • 活跃用户对话成功率高,任务更复杂专业
  • 现有数据集偏向高手用户,不能代表普通人群

尽管越来越多研究关注用户与大语言模型的互动,但多为静态描述,缺乏对个体行为随时间演变的了解。本文分析了约1.2万名随机抽样的微软必应Copilot用户对话轨迹,并与WildChat-4.8M数据集对比。尽管整体呈现显著群体趋势,但个体用户的行为轨迹变化极小,习惯具有高度稳定性。活跃用户对话更成功,执行的任务也更复杂且偏向专业领域。部分趋势在WildChat-4.8M中可见,但该数据集明显偏向高技能“核心”用户。结果表明当前用户行为难以改变,且用户间差异巨大。两数据集对比揭示WildChat无法代表典型人机交互,对后续研究使用提出重要警示。

原文摘要 · Abstract (English)

Although a growing body of research has begun to describe user--LLM interactions, the picture it paints is largely static; little is known about how individual users change their behavior over time. To address this gap, we analyze the conversational trajectories of ~12,000 randomly sampled Microsoft Bing Copilot users and compare these with data from WildChat-4.8M. While the Copilot data contains significant population-level trends, we find that trends in individual user trajectories are much weaker; user habits prove to be overwhelmingly sticky. We also find stark differences between users of different activity levels: more active users have more successful conversations and use the LLM for more complex and professionally oriented tasks. Some user trends also appear in WildChat-4.8M, but we find evidence that this dataset is significantly skewed towards highly proficient "power" users. Ultimately, our results suggest that existing user behavior is difficult to change and demonstrate the extent of user heterogeneity. Our comparison between datasets highlights that WildChat does not represent typical user--AI interactions, an important caveat for downstream uses of the data.

用户行为大模型对话分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。