构建首个混合类型电商对话数据集,解决大模型在复杂规则下的幻觉问题。
Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules
- 基于真实客服对话构建混合类型数据集,含4799条样本。
- 现有模型在复杂规则下易产生幻觉,性能普遍不足。
- 适合研究电商对话、大模型推理与规则约束的学者使用。
电商平台智能体在帮助用户完成购物需求方面发挥重要作用。为推动该领域研究与应用,已有基准框架用于评估大语言模型在电商场景中的表现。然而,当前基准尚未涵盖混合类型电商对话及复杂领域规则的评测。为此,本文首次提出新语料库Mix-ECom,基于真实客户-客服对话并经去隐私化与添加思维链(CoT)处理而成。Mix-ECom包含4,799个样本,每个对话融合多种对话类型(问答、推荐、任务导向、闲聊),覆盖三种电商任务类型(售前、物流、售后)及82条电商规则。同时,本文在该数据集上建立基线,并提出动态优化框架以提升性能。实验表明,当前电商智能体因复杂规则导致幻觉而能力不足。数据集将公开共享。
原文摘要 · Abstract (English)
E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are introduced for evaluating LLM agents in the e-commerce domain. Despite the progress, current benchmarks lack evaluating agents' capability to handle mixed-type e-commerce dialogue and complex domain rules. To address the issue, this work first introduces a novel corpus, termed Mix-ECom, which is constructed based on real-world customer-service dialogues with post-processing to remove user privacy and add CoT process. Specifically, Mix-ECom contains 4,799 samples with multiply dialogue types in each e-commerce dialogue, covering four dialogue types (QA, recommendation, task-oriented dialogue, and chit-chat), three e-commerce task types (pre-sales, logistics, after-sales), and 82 e-commerce rules. Furthermore, this work build baselines on Mix-Ecom and propose a dynamic framework to further improve the performance. Results show that current e-commerce agents lack sufficient capabilities to handle e-commerce dialogues, due to the hallucination cased by complex domain rules. The dataset will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。