用交易数据训练可解释的顾客行为模型,支持跨商家推荐与决策。
Large Behavior Model: A Promptable Digital Twin of the Retail Customer

- 基于用户历史购买和产品信息,统一建模人与环境互动。
- 在多个零售任务中超越主流语言模型,零样本迁移表现优异。
- 适合需要可解释客户模拟的电商、营销与推荐系统研发者。
客户行为建模支撑推荐、营销与决策支持,但现有方法或只重预测准确率而无法解释决策,或模拟用户却缺乏真实行为数据支撑。本文提出大型行为模型(LBM),通过统一的‘人-环境’框架,直接从大规模零售交易数据中学习客户决策机制。客户状态由历史购买生成的行为画像表示,产品上下文通过检索增强生成引入。模型采用持续预训练、监督微调与可验证奖励的强化学习进行训练。在购买预测、难负例区分、购物篮补全、促销响应及跨域优惠券使用等任务上评估,模型在本领域任务中持续优于前沿通用语言模型,并展现出强大的零样本与微调后的跨零售商和决策领域迁移能力。消融实验表明:持续预训练是行为泛化的主要驱动力;检索在训练与推理阶段均应用时效果最佳;强化学习显著提升模型对显式行为证据的依赖,降低对通用语言模型先验的依赖。结果表明,交易历史中的行为知识可通过语言模型有效学习,为顾客数字孪生与行为模拟提供可扩展基础。
原文摘要 · Abstract (English)
Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data. We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environment formulation. Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented generation. The model is trained using continued pre-training on verbalized behavioral data, supervised fine-tuning for decision generation, and reinforcement learning with verifiable rewards for evidence-based calibration. We evaluate the proposed framework on purchase prediction, hard-negative discrimination, basket completion, promotion response, and cross-domain voucher redemption. The model consistently outperforms frontier general-purpose language models on in-domain retail tasks while demonstrating strong zero-shot and fine-tuned transfer across retailers and decision domains. Ablation studies show that continued pre-training is the primary driver of behavioral generalization, retrieval is most effective when applied during both training and inference, and reinforcement learning improves reliance on explicit behavioral evidence over generic language-model priors. These results demonstrate that behavioral knowledge encoded in transaction histories can be effectively learned by language models, providing a scalable foundation for customer digital twins and behavior simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。