arXiv:2507.18884cs.CL2025-07

让客服机器人自我进化,更懂电商对话的复杂需求。

MindFlow+: A Self-Evolving Agent for E-Commerce Customer Service

  • 用大模型+模仿学习+离线强化学习,让机器人从真实对话中持续优化。
  • 在真实电商对话中,上下文相关性、灵活性和任务准确率均超越基线。
  • 引入新指标衡量AI参与度,适合需要高智能客服的电商场景。

高质量对话对电商客服至关重要,但传统基于意图的系统难以应对动态多轮交互。我们提出MindFlow+,一种通过结合大语言模型(LLMs)、模仿学习与离线强化学习(RL)实现自进化的对话代理。MindFlow+引入两种以数据为中心的机制:工具增强的示范构建,使模型接触知识增强且具代理行为(ReAct风格)的交互,提升工具使用能力;以及奖励条件数据建模,利用奖励信号对齐响应与特定任务目标。为评估模型在生成回复中的作用,我们提出全新指标AI贡献率,量化AI在对话中的参与程度。在真实电商对话数据上的实验表明,MindFlow+在上下文相关性、灵活性和任务准确性方面均优于强基线。结果证明,结合大模型、工具推理与奖励引导学习,可构建领域专业化、上下文感知的对话系统。

原文摘要 · Abstract (English)

High-quality dialogue is crucial for e-commerce customer service, yet traditional intent-based systems struggle with dynamic, multi-turn interactions. We present MindFlow+, a self-evolving dialogue agent that learns domain-specific behavior by combining large language models (LLMs) with imitation learning and offline reinforcement learning (RL). MindFlow+ introduces two data-centric mechanisms to guide learning: tool-augmented demonstration construction, which exposes the model to knowledge-enhanced and agentic (ReAct-style) interactions for effective tool use; and reward-conditioned data modeling, which aligns responses with task-specific goals using reward signals. To evaluate the model's role in response generation, we introduce the AI Contribution Ratio, a novel metric quantifying AI involvement in dialogue. Experiments on real-world e-commerce conversations show that MindFlow+ outperforms strong baselines in contextual relevance, flexibility, and task accuracy. These results demonstrate the potential of combining LLMs tool reasoning, and reward-guided learning to build domain-specialized, context-aware dialogue systems.

对话系统大模型电商客服自进化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。