arXiv:2508.16478cs.CLcs.IR2025-08

用大模型做分级文本分类,省资源还易维护

LLM-as-classifier: Semi-Supervised, Iterative Framework for Hierarchical Text Classification using Large Language Models

  • 利用大模型零样本/少样本能力,迭代优化分类框架
  • 通过人机协同和多轮验证,提升分类准确率与可解释性
  • 适合需要持续更新的工业级文本分类场景

大型语言模型(LLMs)为分析非结构化文本数据提供了前所未有的能力。然而,在生产环境中将这些模型部署为可靠、鲁棒且可扩展的分类器,仍面临重大方法论挑战。标准微调方法资源消耗大,且常难以应对现实数据分布的动态变化。本文提出一种全面的半监督框架,利用大模型的零样本和少样本能力,构建用于分级文本分类的解决方案。该方法强调迭代式、人机协同的过程,从领域知识提取开始,经提示优化、层级扩展及多维度验证。我们引入评估并缓解序列偏差的技术,并制定持续监控与适应协议。该框架旨在弥合大模型原始能力与工业应用中对精准、可解释、可维护分类系统的需求之间的差距。

原文摘要 · Abstract (English)

The advent of Large Language Models (LLMs) has provided unprecedented capabilities for analyzing unstructured text data. However, deploying these models as reliable, robust, and scalable classifiers in production environments presents significant methodological challenges. Standard fine-tuning approaches can be resource-intensive and often struggle with the dynamic nature of real-world data distributions, which is common in the industry. In this paper, we propose a comprehensive, semi-supervised framework that leverages the zero- and few-shot capabilities of LLMs for building hierarchical text classifiers as a framework for a solution to these industry-wide challenges. Our methodology emphasizes an iterative, human-in-the-loop process that begins with domain knowledge elicitation and progresses through prompt refinement, hierarchical expansion, and multi-faceted validation. We introduce techniques for assessing and mitigating sequence-based biases and outline a protocol for continuous monitoring and adaptation. This framework is designed to bridge the gap between the raw power of LLMs and the practical need for accurate, interpretable, and maintainable classification systems in industry applications.

大模型分类半监督学习分级文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。