用大模型生成跨领域情感分析数据,解决真实对话数据稀缺问题
Multi-Domain ABSA Conversation Dataset Generation via LLMs for Real-World Evaluation and Model Comparison
- 用GPT-4o生成多领域一致分布的合成对话数据
- 三款主流大模型在合成数据上表现各异:深求R1精度高,谷歌与Claude召回强,谷歌推理快
- 适合需要高质量评估数据的研究者,尤其缺真实标注数据时
基于方面的情感分析(ABSA)能提供细粒度观点洞察,但缺乏反映真实对话细节的多样化标注数据。本文提出一种基于大语言模型(LLMs)的合成ABSA数据生成方法,使用GPT-4o构建跨多个领域的数据,确保主题与情感分布的一致性。通过评估Gemini 1.5 Pro、Claude 3.5 Sonnet和DeepSeek-R1三款先进大模型在主题与情感分类任务上的表现,验证了合成数据的质量与实用性。结果表明:DeepSeek-R1表现出更高精度,Gemini 1.5 Pro与Claude 3.5 Sonnet具有更强召回能力,而Gemini 1.5 Pro推理速度显著更快。研究证明,基于大模型的合成数据生成是创建有价值ABSA资源的有效且灵活的方法,可减少对有限或难以获取的真实标注数据的依赖。
原文摘要 · Abstract (English)
Aspect-Based Sentiment Analysis (ABSA) offers granular insights into opinions but often suffers from the scarcity of diverse, labeled datasets that reflect real-world conversational nuances. This paper presents an approach for generating synthetic ABSA data using Large Language Models (LLMs) to address this gap. We detail the generation process aimed at producing data with consistent topic and sentiment distributions across multiple domains using GPT-4o. The quality and utility of the generated data were evaluated by assessing the performance of three state-of-the-art LLMs (Gemini 1.5 Pro, Claude 3.5 Sonnet, and DeepSeek-R1) on topic and sentiment classification tasks. Our results demonstrate the effectiveness of the synthetic data, revealing distinct performance trade-offs among the models: DeepSeekR1 showed higher precision, Gemini 1.5 Pro and Claude 3.5 Sonnet exhibited strong recall, and Gemini 1.5 Pro offered significantly faster inference. We conclude that LLM-based synthetic data generation is a viable and flexible method for creating valuable ABSA resources, facilitating research and model evaluation without reliance on limited or inaccessible real-world labeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。