构建首个用户在线购物行为模拟数据集,评测大模型预测能力
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation

- 通过问卷与浏览器插件采集真实用户的行为、人格、理由与操作
- 首次建立基于人物画像和实时推理的行动预测基准
- 适合研究个性化数字孪生与智能代理的学者使用
大型语言模型(LLMs)能否准确模拟特定用户的下一步网络操作?尽管LLMs在生成“可信”的人类行为方面展现出潜力,但评估其模仿真实用户行为的能力仍面临挑战,主要源于缺乏高质量、公开可用的数据集,这些数据集需同时捕捉可观测行为与用户内在推理。为此,我们提出OPERA——一个包含观察、人格、理由与行动的新型数据集,来自真实用户在线购物过程中的参与。OPERA是首个公开数据集,全面涵盖用户人格、浏览器观察、细粒度网页操作及即时自述理由。我们开发了在线问卷与定制浏览器插件,以高保真度收集该数据。利用OPERA,我们建立了首个基准,用于评估当前LLMs在给定人格及历史观察-行动-理由序列下,预测特定用户下一步操作与理由的能力。该数据集为未来研究旨在作为个性化数字孪生的人工智能代理奠定了基础。
原文摘要 · Abstract (English)
Can large language models (LLMs) accurately simulate the next web action of a specific user? While LLMs have shown promising capabilities in generating ``believable'' human behaviors, evaluating their ability to mimic real user behaviors remains an open challenge, largely due to the lack of high-quality, publicly available datasets that capture both the observable actions and the internal reasoning of an actual human user. To address this gap, we introduce OPERA, a novel dataset of Observation, Persona, Rationale, and Action collected from real human participants during online shopping sessions. OPERA is the first public dataset that comprehensively captures: user personas, browser observations, fine-grained web actions, and self-reported just-in-time rationales. We developed both an online questionnaire and a custom browser plugin to gather this dataset with high fidelity. Using OPERA, we establish the first benchmark to evaluate how well current LLMs can predict a specific user's next action and rationale with a given persona and <observation, action, rationale> history. This dataset lays the groundwork for future research into LLM agents that aim to act as personalized digital twins for human.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。