arXiv:2503.04830cs.CLcs.AI2025-03被引 13

让电商对话机器人回答时附带引用来源,提升准确性和可信度。

Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents

  • 在回复中自动生成文献引用,增强内容事实依据。
  • 线上测试显示用户参与度提升3%至10%。
  • 支持大规模部署,不改变原有用户体验。

随着对话式大语言模型(LLM)的发展,基于LLM的电商购物助手(CSA)被广泛用于帮助客户流畅购物。构建有吸引力且可信的CSA核心目标是确保其关于产品信息的回复准确且有事实依据。然而仍面临两大挑战:第一,LLM会产生幻觉或无依据的陈述,可能传播错误信息并削弱客户信任;第二,若未在回复中提供知识来源,客户难以验证生成信息的真实性。为此,本文提出一种可快速投入生产的“引用体验”方案。我们构建了自动评估指标,全面评测LLM的依据性与溯源能力,结果显示引用生成范式使依据性表现提升13.83%。为实现规模化部署,我们引入Multi-UX-Inference系统,在保持原有用户体验的前提下,将来源引用附加至LLM输出。大规模在线A/B测试表明,具备事实依据的回复使客户参与度提升3%–10%,具体效果依用户体验设计而异。

原文摘要 · Abstract (English)

With the advancement of conversational large language models (LLMs), several LLM-based Conversational Shopping Agents (CSA) have been developed to help customers smooth their online shopping. The primary objective in building an engaging and trustworthy CSA is to ensure the agent's responses about product factoids are accurate and factually grounded. However, two challenges remain. First, LLMs produce hallucinated or unsupported claims. Such inaccuracies risk spreading misinformation and diminishing customer trust. Second, without providing knowledge source attribution in CSA response, customers struggle to verify LLM-generated information. To address both challenges, we present an easily productionized solution that enables a ''citation experience'' to our customers. We build auto-evaluation metrics to holistically evaluate LLM's grounding and attribution capabilities, suggesting that citation generation paradigm substantially improves grounding performance by 13.83%. To deploy this capability at scale, we introduce Multi-UX-Inference system, which appends source citations to LLM outputs while preserving existing user experience features and supporting scalable inference. Large-scale online A/B tests show that grounded CSA responses improves customer engagement by 3% - 10%, depending on UX variations.

电商对话事实依据引用生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。