用可解释性指导大模型生成情感文本数据,提升小众情绪识别效果
An Interpretability-Guided Framework for Responsible Synthetic Data Generation in Emotional Text
- 基于SHAP的可解释性引导,优化大模型生成情感文本
- 生成数据在小众情绪分类上性能显著提升,接近真实数据
- 适合关注数据伦理与模型可信度的研究者使用
从社交媒体中识别情绪对理解公众情绪至关重要,但因API成本上升和平台限制,获取训练数据变得极为昂贵。本文提出一种可解释性引导框架,利用沙普利加法解释(SHAP)为基于大模型的合成数据生成提供原则性指导。在有足够种子数据的前提下,该方法生成的数据性能接近真实数据,显著优于简单生成方式,并大幅改善了少数情绪类别的分类表现。然而语言学分析显示,合成文本词汇丰富度较低,且缺乏个人化或时间复杂性的表达。本研究既提供了负责任合成数据生成的实用方案,也揭示了其局限性,强调未来可信AI的发展需在合成数据效用与真实感之间权衡。
原文摘要 · Abstract (English)
Emotion recognition from social media is critical for understanding public sentiment, but accessing training data has become prohibitively expensive due to escalating API costs and platform restrictions. We introduce an interpretability-guided framework where Shapley Additive Explanations (SHAP) provide principled guidance for LLM-based synthetic data generation. With sufficient seed data, SHAP-guided approach matches real data performance, significantly outperforms naïve generation, and substantially improves classification for underrepresented emotion classes. However, our linguistic analysis reveals that synthetic text exhibits reduced vocabulary richness and fewer personal or temporally complex expressions than authentic posts. This work provides both a practical framework for responsible synthetic data generation and a critical perspective on its limitations, underscoring that the future of trustworthy AI depends on navigating the trade-offs between synthetic utility and real-world authenticity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。