用大数据技术实时分析社交帖子中的压力情绪
Real-time stress detection on social network posts using big data technology
- 基于Reddit数据构建实时压力检测系统
- 模型在流数据上准确率达69.39%,F1分数达68.97%
- 适合心理健康监测与社交媒体分析研究者
在工业4.0的在线环境中,情绪与心境常通过社交平台帖子表达。这类内容生成了海量且有价值的大数据资源,为自动化、精准的压力检测提供了机遇。本研究构建了一个实时压力检测系统,使用Dreaddit数据集(包含187,444条来自五个Reddit子版块的帖子),其中每类包含压力与非压力文本,体现多样化的压力表达。针对训练,构建了包含3,553条标注数据的语料库。系统采用Apache Kafka、PySpark和AirFlow实现部署。逻辑回归模型在新流数据上表现最佳,准确率为69.39%,F1分数为68.97%。
原文摘要 · Abstract (English)
In the context of modern life, particularly in Industry 4.0 within the online space, emotions and moods are frequently conveyed through social media posts. The trend of sharing stories, thoughts, and feelings on these platforms generates a vast and promising data source for Big Data. This creates both a challenge and an opportunity for research in applying technology to develop more automated and accurate methods for detecting stress in social media users. In this study, we developed a real-time system for stress detection in online posts, using the "Dreaddit: A Reddit Dataset for Stress Analysis in Social Media," which comprises 187,444 posts across five different Reddit domains. Each domain contains texts with both stressful and non-stressful content, showcasing various expressions of stress. A labeled dataset of 3,553 lines was created for training. Apache Kafka, PySpark, and AirFlow were utilized to build and deploy the model. Logistic Regression yielded the best results for new streaming data, achieving 69,39% for measuring accuracy and 68,97 for measuring F1-scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。