arXiv:2412.13098cs.CLcs.SI2024-12

构建1.4万条肯尼亚大选公民报告数据集,助力社会公益领域智能分析。

Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan Election

  • 收集并标注1.4万条选举相关公民报告,含分类与地理标签。
  • 验证语言模型可高效辅助大规模报告的分类与定位。
  • 适合关注AI向善、选举监测与开放数据研究者使用。

在线报告平台使全球公民能够实时分享意见并上报影响本地社区的事件。系统化整理(如按属性分类)和地理标记大量众包信息,对确保从数据中提取准确有意义的洞察至关重要,从而帮助政策制定者推动积极变革。然而,这些任务通常需要大量人工标注。本文提出Uchaguzi-2022,一个包含1.4万条已分类和地理标记的公民报告数据集,涵盖2022年肯尼亚大选中的选举相关议题,如官方渎职、计票异常及暴力行为。我们利用该数据集研究语言模型是否能规模化辅助报告分类与地理标记,凸显其在人工智能助力社会公益领域的潜力。

原文摘要 · Abstract (English)

Online reporting platforms have enabled citizens around the world to collectively share their opinions and report in real time on events impacting their local communities. Systematically organizing (e.g., categorizing by attributes) and geotagging large amounts of crowdsourced information is crucial to ensuring that accurate and meaningful insights can be drawn from this data and used by policy makers to bring about positive change. These tasks, however, typically require extensive manual annotation efforts. In this paper we present Uchaguzi-2022, a dataset of 14k categorized and geotagged citizen reports related to the 2022 Kenyan General Election containing mentions of election-related issues such as official misconduct, vote count irregularities, and acts of violence. We use this dataset to investigate whether language models can assist in scalably categorizing and geotagging reports, thus highlighting its potential application in the AI for Social Good space.

数据集选举监测社会公益语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。