arXiv:2409.08931cs.IR2024-09被引 3

用大模型自动标注视频搜索意图,提升准确率与实时性。

LLM-based Weak Supervision Framework for Query Intent Classification in Video Search

  • 用提示工程和多角色LLM生成标注数据,替代人工标注。
  • 召回率相对基线提升113%,标注一致性提高47.60%。
  • 适合需要快速迭代的视频搜索系统研发团队使用。

流媒体服务重塑了我们发现和参与数字娱乐的方式。尽管如此,有效理解用户搜索查询的广泛语义仍面临重大挑战。一个能处理多种实体代表不同用户意图的精准查询理解系统,对提升用户体验至关重要。可通过训练自然语言理解(NLU)模型实现,但获取该专业领域的高质量标注数据仍是巨大障碍。手动标注成本高且难以覆盖用户词汇的多样性。为此,我们提出一种新颖方法,利用大语言模型(LLMs)通过弱监督自动标注海量用户搜索查询。通过提示工程和多样化的LLM角色,生成符合人类标注者预期的训练数据。结合链式思维(Chain of Thought)和上下文学习(In-Context Learning)融入领域知识,我们的方法利用标注数据训练低延迟模型,适用于实时推理。大量评估表明,本方法在召回率上相较基线平均提升113%。此外,我们的新提示工程框架生成的标注数据质量更高,与人类标注的F1得分一致率提升47.60%(按查询出现频次加权)。人物角色选择路由机制在该框架基础上进一步提升加权F1得分3.67%。

原文摘要 · Abstract (English)

Streaming services have reshaped how we discover and engage with digital entertainment. Despite these advancements, effectively understanding the wide spectrum of user search queries continues to pose a significant challenge. An accurate query understanding system that can handle a variety of entities that represent different user intents is essential for delivering an enhanced user experience. We can build such a system by training a natural language understanding (NLU) model; however, obtaining high-quality labeled training data in this specialized domain is a substantial obstacle. Manual annotation is costly and impractical for capturing users' vast vocabulary variations. To address this, we introduce a novel approach that leverages large language models (LLMs) through weak supervision to automatically annotate a vast collection of user search queries. Using prompt engineering and a diverse set of LLM personas, we generate training data that matches human annotator expectations. By incorporating domain knowledge via Chain of Thought and In-Context Learning, our approach leverages the labeled data to train low-latency models optimized for real-time inference. Extensive evaluations demonstrated that our approach outperformed the baseline with an average relative gain of 113% in recall. Furthermore, our novel prompt engineering framework yields higher quality LLM-generated data to be used for weak supervision; we observed 47.60% improvement over baseline in agreement rate between LLM predictions and human annotations with respect to F1 score, weighted according to the distribution of occurrences of the search queries. Our persona selection routing mechanism further adds an additional 3.67% increase in weighted F1 score on top of our novel prompt engineering framework.

大模型弱监督视频搜索意图识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。