用智能体生成更精准的评论情感分析数据,提升模型效果。
Label-Consistent Data Generation for Aspect-Based Sentiment Analysis Using LLM Agents
- 设计智能体迭代生成并验证数据,保证标签一致性。
- 在三个子任务上,智能体方法比普通提示生成更保真。
- 尤其对T5-Base模型提升明显,适合小模型训练增强。
我们提出一种面向方面情感分析(ABSA)的智能体式数据增强方法,通过迭代生成与验证生成高质量合成训练样本。为隔离智能体结构的影响,还构建了一个使用相同模型和指令的提示基基线。两种方法在三个ABSA子任务(方面词提取ATE、方面情感分类ATSC、方面情感对提取ASPE)、四个SemEval数据集及两种编码器-解码器模型(T5-Base和Tk-Instruct)上进行评估。结果表明,智能体增强在标签保留方面优于原始提示,在需生成方面词的任务中尤为显著。当与真实数据结合时,智能体方法持续优于提示基生成,优势在T5-Base上最明显,而预训练更充分的Tk-Instruct则提升较小。最终,经增强的数据使T5-Base性能接近其对应基准。
原文摘要 · Abstract (English)
We propose an agentic data augmentation method for Aspect-Based Sentiment Analysis (ABSA) that uses iterative generation and verification to produce high quality synthetic training examples. To isolate the effect of agentic structure, we also develop a closely matched prompting-based baseline using the same model and instructions. Both methods are evaluated across three ABSA subtasks (Aspect Term Extraction (ATE), Aspect Sentiment Classification (ATSC), and Aspect Sentiment Pair Extraction (ASPE)), four SemEval datasets, and two encoder-decoder models: T5-Base and Tk-Instruct. Our results show that the agentic augmentation outperforms raw prompting in label preservation of the augmented data, especially when the tasks require aspect term generation. In addition, when combined with real data, agentic augmentation provides higher gains, consistently outperforming prompting-based generation. These benefits are most pronounced for T5-Base, while the more heavily pretrained Tk-Instruct exhibits smaller improvements. As a result, augmented data helps T5-Base achieve comparable performance with its counterpart.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。