arXiv:2602.23716cs.AI2026-02ACL被引 3

用多智能体生成电商深度调研数据,提升AI购物助手的分析能力。

ProductResearch: Training E-Commerce Deep Research Agents via Multi-Agent Synthetic Trajectory Distillation

论文配图:ProductResearch: Training E-Commerce Deep Research Agents via Multi-Agent Synthetic Trajectory Distillation
图 1 · 摘自论文原文
  • 构建用户、研究与监督三智能体协同生成长序列购物轨迹。
  • 微调后模型在信息深度和用户体验上显著超越基线,接近顶级商业系统。
  • 适合需要复杂产品比对与决策支持的电商AI研发团队使用。

基于大语言模型的智能体在电商对话购物中展现潜力,但现有方案缺乏复杂产品研究所需的交互深度与上下文广度。尽管深度研究范式在网页搜索中推进了信息整合,但在电商领域仍存在领域差距。本文提出ProductResearch,一种多智能体框架,通过合成高保真、长时序工具使用轨迹,训练鲁棒的电商购物智能体。该框架包含用户代理(根据行为历史推断细微购物意图)、研究代理与监督代理,后者协调迭代协作生成涵盖全面洞察的产品研究报告。合成轨迹经严格筛选与反思式内化过程,将多智能体监督互动压缩为连贯单角色训练样本,实现对LLM智能体的有效微调。大量实验表明,基于合成数据微调的小型MoE模型在响应完整性、研究深度及用户感知效用方面显著优于基线模型,性能逼近前沿专有深度研究系统,确立多智能体合成轨迹训练是提升基于LLM购物辅助的有效且可扩展范式。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based agents show promise for e-commerce conversational shopping, yet existing implementations lack the interaction depth and contextual breadth required for complex product research. Meanwhile, the Deep Research paradigm, despite advancing information synthesis in web search, suffers from domain gaps when transferred to e-commerce. We propose ProductResearch, a multi-agent framework that synthesizes high-fidelity, long-horizon tool-use trajectories for training robust e-commerce shopping agents. The framework employs a User Agent to infer nuanced shopping intents from behavioral histories, and a Supervisor Agent that orchestrates iterative collaboration with a Research Agent to generate synthetic trajectories culminating in comprehensive, insightful product research reports. These trajectories are rigorously filtered and distilled through a reflective internalization process that consolidates multi-agent supervisory interactions into coherent single-role training examples, enabling effective fine-tuning of LLM agents for complex shopping inquiries. Extensive experiments show that a compact MoE model fine-tuned on our synthetic data achieves substantial improvements over its base model in response comprehensiveness, research depth, and user-perceived utility, approaching the performance of frontier proprietary deep research systems and establishing multi-agent synthetic trajectory training as an effective and scalable paradigm for enhancing LLM-based shopping assistance.

电商AI多智能体深度研究轨迹生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。