arXiv:2510.19245cs.CYcs.AI2025-10被引 9

用视觉+文本双模态模拟网购行为,更接近真人决策。

See, Think, Act: Online Shopper Behavior Simulation with VLM Agents

  • 结合网页截图与文本,让智能体基于多模态信息做判断。
  • 视觉信息使准确率提升6%以上,复杂场景表现更好。
  • 适合研究用户行为建模、人机交互的学者与工程师。

大型语言模型在模拟在线购物行为方面展现出强大潜力。以往工作通过在动作轨迹上进行监督微调(SFT)并引入大模型生成推理过程,以及利用强化学习(RL)提升推理能力,改善了动作预测效果。然而,现有方法依赖纯文本输入,忽视了视觉感知在网页图形界面交互中对人类决策的关键作用。本文探索将网页截图等视觉信息融入基于视觉语言模型(VLM)的行为模拟中,基于OPeRA数据集开展研究。通过同时使用文本和视觉模态作为决策依据,旨在缩小合成智能体与真实用户之间的差距,实现更符合认知规律的购物行为仿真。具体而言,采用联合训练策略,在包含历史动作、过往HTML观测和当前网页截图的完整上下文条件下进行动作预测与推理生成。为进一步增强推理能力,引入分层奖励机制,并设计难度感知因子以优先关注高难度决策点。实验表明,融合视觉信息后,准确率相比仅用文本输入提升了超过6%。结果表明,多模态接地不仅提高了预测精度,还显著增强了在视觉复杂环境下的仿真保真度,捕捉到文本模型常忽略的人类注意力分布与决策细节。最后,本文重新审视行为模拟框架的设计空间,识别关键方法局限,并提出未来构建高效、有效人类行为模拟器的研究方向。

原文摘要 · Abstract (English)

LLMs have recently demonstrated strong potential in simulating online shopper behavior. Prior work has improved action prediction by applying SFT on action traces with LLM-generated rationales, and by leveraging RL to further enhance reasoning capabilities. Despite these advances, current approaches rely on text-based inputs and overlook the essential role of visual perception in shaping human decision-making during web GUI interactions. In this paper, we investigate the integration of visual information, specifically webpage screenshots, into behavior simulation via VLMs, leveraging OPeRA dataset. By grounding agent decision-making in both textual and visual modalities, we aim to narrow the gap between synthetic agents and real-world users, thereby enabling more cognitively aligned simulations of online shopping behavior. Specifically, we employ SFT for joint action prediction and rationale generation, conditioning on the full interaction context, which comprises action history, past HTML observations, and the current webpage screenshot. To further enhance reasoning capabilities, we integrate RL with a hierarchical reward structure, scaled by a difficulty-aware factor that prioritizes challenging decision points. Empirically, our studies show that incorporating visual grounding yields substantial gains: the combination of text and image inputs improves exact match accuracy by more than 6% over text-only inputs. These results indicate that multi-modal grounding not only boosts predictive accuracy but also enhances simulation fidelity in visually complex environments, which captures nuances of human attention and decision-making that text-only agents often miss. Finally, we revisit the design space of behavior simulation frameworks, identify key methodological limitations, and propose future research directions toward building efficient and effective human behavior simulators.

行为模拟视觉语言模型多模态用户研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。