arXiv:2502.13481cs.IR2025-02KDD被引 5

用大模型自动打标签,提升搜索推荐的精准度

LLM4Tag: Automatic Tagging System for Information Retrieval via Large Language Models

  • 构建图结构召回候选标签,确保覆盖全面
  • 融合长短期知识生成准确标签,效果优于现有方法
  • 量化标签置信度,适合大规模在线部署

标签系统在搜索引擎和推荐系统等信息检索场景中至关重要。近期,大语言模型(LLM)因其丰富的世界知识、语义理解与推理能力被用于标签生成。尽管表现优异,现有方法仍存在三方面局限:难以全面召回相关候选标签、难适应新兴领域知识、缺乏可靠的标签置信度评估。为此,我们提出自动标签系统 LLM4Tag。首先设计基于图的标签召回模块,高效构建小规模高相关候选标签集;随后引入知识增强的标签生成模块,通过长期与短期知识注入生成精准标签;最后引入标签置信度校准模块,输出可靠置信分数。在三个大规模工业数据集上的实验表明,LLM4Tag显著优于当前最优基线,已在真实线上环境部署,服务于数亿用户。

原文摘要 · Abstract (English)

Tagging systems play an essential role in various information retrieval applications such as search engines and recommender systems. Recently, Large Language Models (LLMs) have been applied in tagging systems due to their extensive world knowledge, semantic understanding, and reasoning capabilities. Despite achieving remarkable performance, existing methods still have limitations, including difficulties in retrieving relevant candidate tags comprehensively, challenges in adapting to emerging domain-specific knowledge, and the lack of reliable tag confidence quantification. To address these three limitations above, we propose an automatic tagging system LLM4Tag. First, a graph-based tag recall module is designed to effectively and comprehensively construct a small-scale highly relevant candidate tag set. Subsequently, a knowledge-enhanced tag generation module is employed to generate accurate tags with long-term and short-term knowledge injection. Finally, a tag confidence calibration module is introduced to generate reliable tag confidence scores. Extensive experiments over three large-scale industrial datasets show that LLM4Tag significantly outperforms the state-of-the-art baselines and LLM4Tag has been deployed online for content tagging to serve hundreds of millions of users.

标签生成大模型应用信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。