构建印尼话语多标签数据集,揭示毒性与极化相互增强关系
A Multi-Labeled Dataset for Indonesian Discourse: Examining Toxicity, Polarization, and Demographics Information
- 构建含毒性、极化与标注者人口信息的多标签印尼语数据集
- 极化信息可提升毒性识别准确率,反之亦然,二者存在协同效应
- 加入标注者背景信息显著改善极化分类效果,适合社会计算研究者
极化指在实质性议题上多个群体持有的对立观点。作为世界第三大民主国家,印度尼西亚日益面临政治极化与网络暴力交织的挑战,后者常针对弱势少数群体。尽管该问题至关重要,但以往自然语言处理研究尚未充分探索毒性与极化之间的关联。为此,本文提出一个全新的多标签印尼语数据集,包含毒性、极化以及标注者人口统计信息。使用BERT-base模型和大型语言模型(LLMs)对数据集进行基准测试发现,极化信息能提升毒性分类性能,反之亦然;同时,提供标注者人口信息显著改善了极化分类表现。
原文摘要 · Abstract (English)
Polarization is defined as divisive opinions held by two or more groups on substantive issues. As the world's third-largest democracy, Indonesia faces growing concerns about the interplay between political polarization and online toxicity, which is often directed at vulnerable minority groups. Despite the importance of this issue, previous NLP research has not fully explored the relationship between toxicity and polarization. To bridge this gap, we present a novel multi-label Indonesian dataset that incorporates toxicity, polarization, and annotator demographic information. Benchmarking this dataset using BERT-base models and large language models (LLMs) shows that polarization information enhances toxicity classification, and vice versa. Furthermore, providing demographic information significantly improves the performance of polarization classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。