用AI自动分类多语言选举报告,提升众包监督的可扩展性。
Scaling Crowdsourced Election Monitoring: Construction and Evaluation of Classification Models for Multilingual and Cross-Domain Classification Settings
- 分两步识别报告有用性并分类信息类型,支持多语言跨域
- 有用性检测F1达77%,信息分类F1达75%,跨域零样本仍达59%
- 适合需处理多语言选举数据的公益组织与研究者
众包选举监督正成为传统方式的补充,但依赖志愿者手动处理报告导致扩展困难。本文提出一种两阶段自动化分类方法:先识别报告是否具有信息量,再将其归类为不同信息类型。实验采用XLM-RoBERTa和SBERT等多语言模型,并融合语言学特征,在多语言、跨领域设置下测试。结果表明,信息量检测的F1分数为77%,信息类型分类为75%。在跨域场景中,零样本设置下F1为59%,少样本设置下达63%。分析显示,模型对英文报告的识别优于斯瓦希里语,可能因训练数据不均衡,提示部署时需谨慎。
原文摘要 · Abstract (English)
The adoption of crowdsourced election monitoring as a complementary alternative to traditional election monitoring is on the rise. Yet, its reliance on digital response volunteers to manually process incoming election reports poses a significant scaling bottleneck. In this paper, we address the challenge of scaling crowdsourced election monitoring by advancing the task of automated classification of crowdsourced election reports to multilingual and cross-domain classification settings. We propose a two-step classification approach of first identifying informative reports and then categorising them into distinct information types. We conduct classification experiments using multilingual transformer models such as XLM-RoBERTa and multilingual embeddings such as SBERT, augmented with linguistically motivated features. Our approach achieves F1-Scores of 77\% for informativeness detection and 75\% for information type classification. We conduct cross-domain experiments, applying models trained in a source electoral domain to a new target electoral domain in zero-shot and few-shot classification settings. Our results show promising potential for model transfer across electoral domains, with F1-Scores of 59\% in zero-shot and 63\% in few-shot settings. However, our analysis also reveals a performance bias in detecting informative English reports over Swahili, likely due to imbalances in the training data, indicating a need for caution when deploying classification models in real-world election scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。