arXiv:2505.24251cs.CLcs.IR2025-05ACL被引 6

通过双阶段框架提升工业级搜索对话的主动引导能力

Proactive Guidance of Multi-Turn Conversation in Industrial Search

  • 分两阶段:先动态适应用户目标,再基于点击信号优化交互
  • 线上点击率提升149%,离线目标识别准确率提高23.95%
  • 适合追求高效、低延迟对话系统的工业应用

大型语言模型的发展推动了多轮对话系统的进步,亟需主动引导以提升用户体验。然而,现有系统在动态适应用户目标变化和保持低延迟方面面临挑战。在百度搜索AI助手这一工业级多轮搜索系统中,我们提出一种新型双阶段框架实现主动引导。第一阶段目标自适应监督微调(G-SFT)利用目标自适应代理动态响应用户目标变化,并提供相关上下文信息;同时通过可扩展的知识蒸馏,将大模型知识迁移到轻量模型以支持实时交互。第二阶段点击导向强化学习(C-RL)采用生成-排序范式,从用户点击信号构建偏好对,通过主动引导提升点击率。该双阶段架构互补:G-SFT确保目标追踪准确性,C-RL则通过点击驱动的强化学习优化交互质量。大量实验表明,该框架在离线评估中达到86.10%的准确率(较基线提升23.95%),在线部署中点击率提升至25.28%(相对提升149.06%),并借助知识蒸馏将推理延迟降低69.55%。

原文摘要 · Abstract (English)

The evolution of Large Language Models (LLMs) has significantly advanced multi-turn conversation systems, emphasizing the need for proactive guidance to enhance users' interactions. However, these systems face challenges in dynamically adapting to shifts in users' goals and maintaining low latency for real-time interactions. In the Baidu Search AI assistant, an industrial-scale multi-turn search system, we propose a novel two-phase framework to provide proactive guidance. The first phase, Goal-adaptive Supervised Fine-Tuning (G-SFT), employs a goal adaptation agent that dynamically adapts to user goal shifts and provides goal-relevant contextual information. G-SFT also incorporates scalable knowledge transfer to distill insights from LLMs into a lightweight model for real-time interaction. The second phase, Click-oriented Reinforcement Learning (C-RL), adopts a generate-rank paradigm, systematically constructs preference pairs from user click signals, and proactively improves click-through rates through more engaging guidance. This dual-phase architecture achieves complementary objectives: G-SFT ensures accurate goal tracking, while C-RL optimizes interaction quality through click signal-driven reinforcement learning. Extensive experiments demonstrate that our framework achieves 86.10% accuracy in offline evaluation (+23.95% over baseline) and 25.28% CTR in online deployment (149.06% relative improvement), while reducing inference latency by 69.55% through scalable knowledge distillation.

对话系统搜索AI强化学习知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。