arXiv:2504.14105cs.HCcs.AI2025-04被引 3

用本地专家共建多语言数据集,提升AI的本地适用性与安全性

Amplify Initiative: Building A Localized Data Platform for Globalized AI

  • 联合本地专家通过移动端应用协同创建数据
  • 生成8091条多语言对抗性问题,覆盖7种非洲语言
  • 适合关注AI本地化、文化敏感性的研究者与开发者

当前AI模型因训练数据以英文和西方网络内容为主,难以反映本地语境与语言,影响其全球适用性、实用性与安全性。Amplify Initiative通过构建数据平台与方法论,联合专家社区收集高质量、多样化的数据以弥补这一缺陷。该平台支持数据共创,提供多语言高质量数据集,并赋予数据贡献者认可。本文展示了在撒哈拉以南非洲五国(加纳、肯尼亚、马拉维、尼日利亚、乌干达)开展的试点项目,与当地研究人员合作,共汇聚155名领域专家(如医生、银行家、人类与公民权利倡导者),利用安卓应用完成端到端数据共创。最终形成包含8,091条对抗性查询的标注数据集,覆盖卢干达语、斯瓦希里语、奇切瓦语等七种语言,涵盖错误信息、公共利益等关键主题的细微语境信息。该数据集可用于评估模型在相关语言中的安全性和文化适配性。

原文摘要 · Abstract (English)

Current AI models often fail to account for local context and language, given the predominance of English and Western internet content in their training data. This hinders the global relevance, usefulness, and safety of these models as they gain more users around the globe. Amplify Initiative, a data platform and methodology, leverages expert communities to collect diverse, high-quality data to address the limitations of these models. The platform is designed to enable co-creation of datasets, provide access to high-quality multilingual datasets, and offer recognition to data authors. This paper presents the approach to co-creating datasets with domain experts (e.g., health workers, teachers) through a pilot conducted in Sub-Saharan Africa (Ghana, Kenya, Malawi, Nigeria, and Uganda). In partnership with local researchers situated in these countries, the pilot demonstrated an end-to-end approach to co-creating data with 155 experts in sensitive domains (e.g., physicians, bankers, anthropologists, human and civil rights advocates). This approach, implemented with an Android app, resulted in an annotated dataset of 8,091 adversarial queries in seven languages (e.g., Luganda, Swahili, Chichewa), capturing nuanced and contextual information related to key themes such as misinformation and public interest topics. This dataset in turn can be used to evaluate models for their safety and cultural relevance within the context of these languages.

数据平台AI本地化多语言协同共创

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。