arXiv:2608.11604cs.AI2026-08

让购物助手直接学用户对话反馈,提升推荐和回复质量。

Learning from Online User Feedback for Shopping Agents

论文配图:Learning from Online User Feedback for Shopping Agents
图 1 · 摘自论文原文
  • 结合购买结果强化学习与对话指令感知的策略蒸馏
  • 在真实电商日志上显著提升推荐准确率与用户满意度
  • 适合想用真实用户交互优化AI导购系统的团队

基于大语言模型的购物助手在真实电商场景中广泛应用,产生了海量用户交互日志,蕴含丰富的监督信号。然而,现有方法多依赖离线训练信号(如用户-商品交互或合成偏好数据),忽视了用户自然对话反馈中的宝贵信息。在线反馈具有异构性、稀疏性和噪声等特点,难以自动转化为可靠学习信号。为此,我们提出LOFA框架,使购物助手可直接从真实在线交互日志中学习,无需人工标注。该框架融合基于可验证购买结果的强化学习与反馈感知的在线策略蒸馏,识别用户对话中的指令并转化为密集的词粒度监督信号。两类目标协同捕捉群体行为模式与用户个性化偏好。在真实电商日志上的大量实验表明,相较于强基线模型,LOFA持续提升了推荐质量、回复帮助性及用户满意度对齐程度,证明了从真实用户反馈中学习的有效性。

原文摘要 · Abstract (English)

Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user interaction logs that provide valuable supervision for improving these agents. However, existing approaches primarily rely on offline training signals, such as user-item interactions or synthetic preference data, while largely overlooking the rich supervision contained in users' natural conversational feedback. Moreover, the available online feedback is heterogeneous, sparse, and noisy, making it difficult to transform into reliable learning signals automatically. To address these challenges, we propose LOFA, a framework that enables shopping agents to learn directly from real online interaction logs without human annotation. LOFA combines reinforcement learning over verifiable purchase outcomes with feedback-aware on-policy distillation, which identifies users'in-dialogue directives and converts them into dense token-level supervision. These complementary objectives capture both collaborative behavioral patterns and user-specific preferences. Extensive experiments on real-world e-commerce logs demonstrate that LOFA consistently improves recommendation quality, response helpfulness, and user-satisfaction alignment over strong baselines, highlighting the effectiveness of learning shopping agents from real online user feedback.

购物代理强化学习对话系统用户反馈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。