用12个数据集构建新文本政治倾向与政治性分类基准
Political Leaning and Politicalness Classification of Texts
- 整合12个数据集,构建更全面的政治倾向与政治性分类数据集
- 通过留一训练/测试验证,发现现有模型在跨数据集上泛化能力差
- 提出新方法提升模型对未见文本的适应性,适合多场景部署
本文针对使用Transformer模型自动分类文本政治倾向与政治性的挑战展开研究。系统梳理了现有数据集和模型,发现当前方法形成信息孤岛,在分布外文本上表现不佳。为此,我们通过整合12个政治倾向分类数据集,并扩展18个现有数据集以新增政治性标签,构建了一个多样化的综合数据集。采用留一入和留一出的基准测试方法,评估了现有模型性能,并训练了具备更强泛化能力的新模型。
原文摘要 · Abstract (English)
This paper addresses the challenge of automatically classifying text according to political leaning and politicalness using transformer models. We compose a comprehensive overview of existing datasets and models for these tasks, finding that current approaches create siloed solutions that perform poorly on out-of-distribution texts. To address this limitation, we compile a diverse dataset by combining 12 datasets for political leaning classification and creating a new dataset for politicalness by extending 18 existing datasets with the appropriate label. Through extensive benchmarking with leave-one-in and leave-one-out methodologies, we evaluate the performance of existing models and train new ones with enhanced generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。