分析9969条中东冲突帖文,比较不同模型识别立场效果
Social media polarization during conflict: Insights from an ideological stance dataset on Israel-Palestine Reddit comments
- 用多种模型分析社交媒体中的政治立场
- 混合专家模型在提示工程下准确率达最高
- 数据集公开,适合研究网络极化现象
在战争等敏感政治情境中,社交媒体常成为观点极化和强烈立场表达的场所。尽管已有研究关注一般语境下的立场识别,但针对冲突场景的研究仍有限。本研究基于2023年10月至2024年8月收集的9,969条关于以巴冲突的Reddit评论,将其分为亲以、亲巴和中立三类立场。采用机器学习、预训练语言模型、神经网络及开源大模型的提示工程策略进行分类,评估指标包括准确率、精确率、召回率和F1分数。结果显示,Mixtral 8x7B结合评分与反思重读提示的方法在所有指标上表现最优。该研究为高极化社交语境下立场检测提供了方法对比,所用数据集已公开,可供进一步探索与验证。
原文摘要 · Abstract (English)
In politically sensitive scenarios like wars, social media serves as a platform for polarized discourse and expressions of strong ideological stances. While prior studies have explored ideological stance detection in general contexts, limited attention has been given to conflict-specific settings. This study addresses this gap by analyzing 9,969 Reddit comments related to the Israel-Palestine conflict, collected between October 2023 and August 2024. The comments were categorized into three stance classes: Pro-Israel, Pro-Palestine, and Neutral. Various approaches, including machine learning, pre-trained language models, neural networks, and prompt engineering strategies for open source large language models (LLMs), were employed to classify these stances. Performance was assessed using metrics such as accuracy, precision, recall, and F1-score. Among the tested methods, the Scoring and Reflective Re-read prompt in Mixtral 8x7B demonstrated the highest performance across all metrics. This study provides comparative insights into the effectiveness of different models for detecting ideological stances in highly polarized social media contexts. The dataset used in this research is publicly available for further exploration and validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。