用混合模型提升社交媒体中性少数群体压力的识别准确率。
Predictive Insights into LGBTQ+ Minority Stress: A Transductive Exploration of Social Media Discourse
- 结合图神经网络与BERT,捕捉语言细微差别。
- 在5789条匿名帖子上达到86%准确率和F1值。
- 适合研究数字健康与社会心理干预的学者。
性少数群体(如女同、男同、双性恋、跨性别等)相比顺性别和异性恋者面临更差的健康状况,主要源于少数群体压力(即适应主流文化时特有的慢性社会压力)。这种压力常通过社交媒体发帖表达,但其语言具有复杂性(如隐喻、词汇多样性),传统自然语言处理方法难以识别。本文设计了一种融合图神经网络(GNN)与预训练模型RoBERTa的混合模型,基于公开的LGBTQ+ MiSSoM+数据集(含5,789条来自性少数社群Reddit帖子的人工标注样本)进行实验。该模型利用大规模原始数据预训练提取潜在语义特征,并通过归纳学习联合优化有标签训练数据与无标签测试数据的表示。RoBERTa-GCN模型在预测中取得0.86的准确率与0.86的F1分数,优于多个基线模型。提升对社交媒体中少数群体压力的识别能力,有助于开发针对性数字健康干预措施,改善该群体高压力相关健康问题。
原文摘要 · Abstract (English)
Individuals who identify as sexual and gender minorities, including lesbian, gay, bisexual, transgender, queer, and others (LGBTQ+) are more likely to experience poorer health than their heterosexual and cisgender counterparts. One primary source that drives these health disparities is minority stress (i.e., chronic and social stressors unique to LGBTQ+ communities' experiences adapting to the dominant culture). This stress is frequently expressed in LGBTQ+ users' posts on social media platforms. However, these expressions are not just straightforward manifestations of minority stress. They involve linguistic complexity (e.g., idiom or lexical diversity), rendering them challenging for many traditional natural language processing methods to detect. In this work, we designed a hybrid model using Graph Neural Networks (GNN) and Bidirectional Encoder Representations from Transformers (BERT), a pre-trained deep language model to improve the classification performance of minority stress detection. We experimented with our model on a benchmark social media dataset for minority stress detection (LGBTQ+ MiSSoM+). The dataset is comprised of 5,789 human-annotated Reddit posts from LGBTQ+ subreddits. Our approach enables the extraction of hidden linguistic nuances through pretraining on a vast amount of raw data, while also engaging in transductive learning to jointly develop representations for both labeled training data and unlabeled test data. The RoBERTa-GCN model achieved an accuracy of 0.86 and an F1 score of 0.86, surpassing the performance of other baseline models in predicting LGBTQ+ minority stress. Improved prediction of minority stress expressions on social media could lead to digital health interventions to improve the wellbeing of LGBTQ+ people-a community with high rates of stress-sensitive health problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。