构建跨领域社交媒体论点立场分类基准,免人工标注。
A Benchmark for Cross-Domain Argumentative Stance Classification on Social Media
- 用平台规则与大模型生成跨领域论点数据
- 覆盖21个领域,含4498个议题与30961条论据
- 支持零样本与少样本学习测试,适合立场分析研究
论点立场分类在识别作者对特定话题观点中起关键作用。然而,跨多个领域生成多样化的论点句子对极具挑战性。现有基准多源于单一领域或仅涵盖有限话题,且人工标注耗时费力。为此,我们提出利用平台规则、现成专家内容及大语言模型,绕过人工标注需求。该方法构建了一个多领域基准,包含4,498个主题性主张和30,961条论据,来自三个来源,覆盖21个领域。我们在全监督、零样本和少样本设置下对该数据集进行基准测试,揭示了不同方法的优劣。本研究已发布数据集与代码,具体信息见隐藏以匿名。
原文摘要 · Abstract (English)
Argumentative stance classification plays a key role in identifying authors' viewpoints on specific topics. However, generating diverse pairs of argumentative sentences across various domains is challenging. Existing benchmarks often come from a single domain or focus on a limited set of topics. Additionally, manual annotation for accurate labeling is time-consuming and labor-intensive. To address these challenges, we propose leveraging platform rules, readily available expert-curated content, and large language models to bypass the need for human annotation. Our approach produces a multidomain benchmark comprising 4,498 topical claims and 30,961 arguments from three sources, spanning 21 domains. We benchmark the dataset in fully supervised, zero-shot, and few-shot settings, shedding light on the strengths and limitations of different methodologies. We release the dataset and code in this study at hidden for anonymity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。