构建多语言暴力事件数据集,助力人道主义援助安全预警
HumVI: A Multilingual Dataset for Detecting Violent Incidents Impacting Humanitarian Aid
- 基于英法阿三语新闻构建暴力事件数据集,按影响领域分类
- 与人道组织合作获取可靠标签,支持安全、教育等五大领域分析
- 提供多种深度学习基准,适用于跨语言、小样本场景
人道主义组织可通过数据分析发现趋势、评估安全风险、支持决策和争取资助。然而,直接关联人道援助的暴力事件数据难以获取。本文提出HumVI——一个涵盖英语、法语、阿拉伯语新闻文章的多语言数据集,包含影响援助安全、教育、粮食安全、健康与保护等领域的暴力事件实例。标签由数据驱动的人道组织Insecurity Insight验证,确保可靠性。我们为该数据集提供了多个基准测试,采用深度学习模型与数据增强、掩码损失等技术应对领域扩展等挑战。数据集已公开于https://github.com/dataminr-ai/humvi-dataset。
原文摘要 · Abstract (English)
Humanitarian organizations can enhance their effectiveness by analyzing data to discover trends, gather aggregated insights, manage their security risks, support decision-making, and inform advocacy and funding proposals. However, data about violent incidents with direct impact and relevance for humanitarian aid operations is not readily available. An automatic data collection and NLP-backed classification framework aligned with humanitarian perspectives can help bridge this gap. In this paper, we present HumVI - a dataset comprising news articles in three languages (English, French, Arabic) containing instances of different types of violent incidents categorized by the humanitarian sector they impact, e.g., aid security, education, food security, health, and protection. Reliable labels were obtained for the dataset by partnering with a data-backed humanitarian organization, Insecurity Insight. We provide multiple benchmarks for the dataset, employing various deep learning architectures and techniques, including data augmentation and mask loss, to address different task-related challenges, e.g., domain expansion. The dataset is publicly available at https://github.com/dataminr-ai/humvi-dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。