构建零样本对话立场检测数据集,提升模型对未知话题的判断能力
Zero-Shot Conversational Stance Detection: Dataset and Approaches
- 提出基于原型对比学习的SITPCL模型,融合说话人互动与目标感知机制
- 在280个未见目标上实现43.81%的F1宏平均分数,达当前最优
- 适合研究零样本推理、社交舆论分析的学者和工程师参考
立场检测旨在通过社交媒体数据识别公众对特定目标的态度,是重要但具有挑战性的任务。随着社交媒体用户间辩论日益增多,对话式立场检测成为关键研究方向。然而,现有对话立场检测数据集仅涵盖有限的具体目标,导致模型在真实场景面对大量未见目标时表现受限。为弥补这一差距,我们人工构建了一个大规模、高质量的零样本对话立场检测数据集ZS-CSD,包含280个目标,覆盖两种不同目标类型。基于该数据集,我们提出SITPCL模型——一种结合说话人交互与目标感知的原型对比学习方法,并建立了零样本设置下的基准性能。实验表明,所提SITPCL模型在零样本对话立场检测中达到领先水平,但其F1-macro得分仅为43.81%,凸显该任务仍面临持续挑战。
原文摘要 · Abstract (English)
Stance detection, which aims to identify public opinion towards specific targets using social media data, is an important yet challenging task. With the increasing number of online debates among social media users, conversational stance detection has become a crucial research area. However, existing conversational stance detection datasets are restricted to a limited set of specific targets, which constrains the effectiveness of stance detection models when encountering a large number of unseen targets in real-world applications. To bridge this gap, we manually curate a large-scale, high-quality zero-shot conversational stance detection dataset, named ZS-CSD, comprising 280 targets across two distinct target types. Leveraging the ZS-CSD dataset, we propose SITPCL, a speaker interaction and target-aware prototypical contrastive learning model, and establish the benchmark performance in the zero-shot setting. Experimental results demonstrate that our proposed SITPCL model achieves state-of-the-art performance in zero-shot conversational stance detection. Notably, the SITPCL model attains only an F1-macro score of 43.81%, highlighting the persistent challenges in zero-shot conversational stance detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。