用双路生成+标签优化,解决少样本情感分析数据不足问题
DS$^2$-ABSA: Dual-Stream Data Synthesis with Label Refinement for Few-Shot Aspect-Based Sentiment Analysis
- 从关键点和实例双视角生成数据,提升多样性
- 合成数据使模型性能超越现有方法,在少样本下表现更优
- 适合低资源场景下的情感分析研究与应用
近期发展的大语言模型(LLMs)为低资源场景下的数据稀缺问题提供了新思路。在少样本方面情感分析(ABSA)中,已有研究尝试通过修改已有样本提示LLM生成新数据,但生成数据多样性不足,效果受限。此外,部分工作采用上下文学习,通过特定指令和少量示例作为提示进行ABSA,尽管有潜力,但LLM常生成不符合任务要求的标签。为克服上述局限,我们提出DS²-ABSA,一种面向少样本ABSA的双流数据合成框架。该框架利用LLM从两个互补视角——关键点驱动与实例驱动——生成多样化且高质量的ABSA样本。同时引入标签精炼模块,进一步提升合成标签的准确性。大量实验表明,DS²-ABSA显著优于现有少样本ABSA方法及其他基于LLM的数据生成方法。
原文摘要 · Abstract (English)
Recently developed large language models (LLMs) have presented promising new avenues to address data scarcity in low-resource scenarios. In few-shot aspect-based sentiment analysis (ABSA), previous efforts have explored data augmentation techniques, which prompt LLMs to generate new samples by modifying existing ones. However, these methods fail to produce adequately diverse data, impairing their effectiveness. Besides, some studies apply in-context learning for ABSA by using specific instructions and a few selected examples as prompts. Though promising, LLMs often yield labels that deviate from task requirements. To overcome these limitations, we propose DS$^2$-ABSA, a dual-stream data synthesis framework targeted for few-shot ABSA. It leverages LLMs to synthesize data from two complementary perspectives: \textit{key-point-driven} and \textit{instance-driven}, which effectively generate diverse and high-quality ABSA samples in low-resource settings. Furthermore, a \textit{label refinement} module is integrated to improve the synthetic labels. Extensive experiments demonstrate that DS$^2$-ABSA significantly outperforms previous few-shot ABSA solutions and other LLM-oriented data generation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。