用大模型检测社交媒体中的社会支持类型,提升识别准确率。
Advanced Machine Learning Techniques for Social Support Detection on Social Media
- 结合聚类与大模型,解决支持内容分类数据不平衡问题。
- 在个体/群体支持识别上提升0.4%、支持类型分类提升0.7%准确率。
- 适合研究网络心理支持、社会情绪分析的学者和从业者。
社交媒体的广泛使用凸显了理解其影响的重要性,尤其是在线社会支持的作用。本研究采用聚焦于在线社会支持的数据集,包含支持内容的二分类与多分类任务。分类任务分为三类:第一类区分支持性与非支持性内容;第二类判断支持对象是个人还是群体;第三类将支持类型归为国家、LGBTQ、黑人、女性、宗教及其他(不属前述类别)。为应对数据不平衡,我们采用K-means聚类进行数据平衡,并与原始未平衡数据对比。通过先进机器学习技术,包括基于Transformer的模型及GPT3、GPT4、GPT4-o的零样本学习方法,预测不同情境下的社会支持水平。实验表明,基于Transformer的方法表现更优。相较以往使用心理语言学与词袋(TF-IDF)的传统机器学习方法,本研究在第二项任务中宏F1得分提升0.4%,第三项任务提升0.7%。
原文摘要 · Abstract (English)
The widespread use of social media highlights the need to understand its impact, particularly the role of online social support. This study uses a dataset focused on online social support, which includes binary and multiclass classifications of social support content on social media. The classification of social support is divided into three tasks. The first task focuses on distinguishing between supportive and non-supportive. The second task aims to identify whether the support is directed toward an individual or a group. The third task categorizes the specific type of social support, grouping it into categories such as Nation, LGBTQ, Black people, Women, Religion, and Other (if it does not fit into the previously mentioned categories). To address data imbalances in these tasks, we employed K-means clustering for balancing the dataset and compared the results with the original unbalanced data. Using advanced machine learning techniques, including transformers and zero-shot learning approaches with GPT3, GPT4, and GPT4-o, we predict social support levels in various contexts. The effectiveness of the dataset is evaluated using baseline models across different learning approaches, with transformer-based methods demonstrating superior performance. Additionally, we achieved a 0.4\% increase in the macro F1 score for the second task and a 0.7\% increase for the third task, compared to previous work utilizing traditional machine learning with psycholinguistic and unigram-based TF-IDF values.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。