用大模型提升文本分类中的数据多样性与类别区分度
TARDiS : Text Augmentation for Refining Diversity and Separability
- 分两阶段生成:用多类别提示增强数据多样性
- 引入类适应机制,确保生成文本准确对齐目标类别
- 适合少样本场景下提升文本分类性能的研究者
文本增强(TA)是文本分类,尤其是少样本设置下的关键技术。本文提出一种基于大模型的新型TA方法TARDiS,解决两阶段TA方法中生成与对齐阶段的固有问题。生成阶段设计SEG和CEG两种生成过程,结合多个类别特定提示,提升数据多样性和类别可区分性;对齐阶段引入类适应(CA)方法,通过验证与修改确保生成样本与目标类别一致。实验表明,TARDiS在多种少样本文本分类任务中优于现有最先进方法。深入分析验证了各阶段的具体行为。
原文摘要 · Abstract (English)
Text augmentation (TA) is a critical technique for text classification, especially in few-shot settings. This paper introduces a novel LLM-based TA method, TARDiS, to address challenges inherent in the generation and alignment stages of two-stage TA methods. For the generation stage, we propose two generation processes, SEG and CEG, incorporating multiple class-specific prompts to enhance diversity and separability. For the alignment stage, we introduce a class adaptation (CA) method to ensure that generated examples align with their target classes through verification and modification. Experimental results demonstrate TARDiS's effectiveness, outperforming state-of-the-art LLM-based TA methods in various few-shot text classification tasks. An in-depth analysis confirms the detailed behaviors at each stage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。