融合符号知识与神经网络,动态捕捉疫情期社交媒体心理舆情。
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
- 结合神经网络与符号知识库,动态适应语言演化。
- 在120亿推文数据上F1超92%,优于纯数据驱动模型。
- 响应快、算力低,适合实时健康监测场景。
通过社交媒体监测公众情绪在新冠疫情等健康危机中具有潜在价值。但传统基于频率和数据驱动的神经网络方法,因语言持续演变而可能遗漏新出现的相关内容。人工构建的符号知识源(如标准语和俚语词典)可提升对动态语言中社交信号的识别能力。本文提出一种神经符号方法,将神经网络与符号知识源融合,增强对与新冠肺炎相关心理健康推文的检测与解读。该方法在约120亿条推文、250万条子版块数据及70万篇新闻文章组成的大型语料库上,结合多个知识图谱进行评估。结果表明,该方法能动态适应语言变化,在多项指标上超越纯数据驱动模型,F1得分超过92%。同时,其对新数据的适应速度更快,计算开销远低于微调大型语言模型(LLMs)。研究证明神经符号方法在动态环境中解释文本(如公共卫生监测)中的有效性。
原文摘要 · Abstract (English)
Monitoring public sentiment via social media is potentially helpful during health crises such as the COVID-19 pandemic. However, traditional frequency-based, data-driven neural network-based approaches can miss newly relevant content due to the evolving nature of language in a dynamically evolving environment. Human-curated symbolic knowledge sources, such as lexicons for standard language and slang terms, can potentially elevate social media signals in evolving language. We introduce a neurosymbolic method that integrates neural networks with symbolic knowledge sources, enhancing the detection and interpretation of mental health-related tweets relevant to COVID-19. Our method was evaluated using a corpus of large datasets (approximately 12 billion tweets, 2.5 million subreddit data, and 700k news articles) and multiple knowledge graphs. This method dynamically adapts to evolving language, outperforming purely data-driven models with an F1 score exceeding 92\%. This approach also showed faster adaptation to new data and lower computational demands than fine-tuning pre-trained large language models (LLMs). This study demonstrates the benefit of neurosymbolic methods in interpreting text in a dynamic environment for tasks such as health surveillance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。