arXiv:2504.21398cs.IR2025-04中稿 · International ACM …被引 4

对比弱监督与大模型在短查询意图分类中的表现

In a Few Words: Comparing Weak Supervision and LLMs for Short Query Intent Classification

  • 用弱监督和大模型分别进行短查询意图分类实验
  • 大模型召回率更高,但精确率仍不理想
  • 适合关注意图识别精度与实用性的研究者

用户意图分类是信息检索中的关键任务。过去,意图分类依赖人工标注或自动方法,后者避免了大规模人工标注。近年来,研究探索大模型是否能可靠判断用户意图,但学界已意识到生成式大模型在分类任务中的局限性。本研究通过实证比较弱监督与大模型在信息型、导航型、交易型三类意图分类中的表现,评估了LLaMA-3.1-8B-Instruct与LLaMA-3.1-70B-Instruct的上下文学习能力,以及前者微调后的效果,并与基于弱监督训练的基准分类器ORCAS-I进行对比。结果表明,尽管大模型在召回率上优于弱监督,但在精确率方面仍存在明显不足,凸显了在两者间实现有效平衡的必要性。

原文摘要 · Abstract (English)

User intent classification is an important task in information retrieval. Previously, user intents were classified manually and automatically; the latter helped to avoid hand labelling of large datasets. Recent studies explored whether LLMs can reliably determine user intent. However, researchers have recognized the limitations of using generative LLMs for classification tasks. In this study, we empirically compare user intent classification into informational, navigational, and transactional categories, using weak supervision and LLMs. Specifically, we evaluate LLaMA-3.1-8B-Instruct and LLaMA-3.1-70B-Instruct for in-context learning and LLaMA-3.1-8B-Instruct for fine-tuning, comparing their performance to an established baseline classifier trained using weak supervision (ORCAS-I). Our results indicate that while LLMs outperform weak supervision in recall, they continue to struggle with precision, which shows the need for improved methods to balance both metrics effectively.

意图分类大模型弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。