arXiv:2409.05247cs.CL2024-09被引 19

为少资源语言构建伦理数据,避免数字殖民

Socially Responsible Data for Large Multilingual Language Models

  • 通过社区合作与参与式设计收集多语言数据
  • 提出12条针对非西方语言数据采集的伦理建议
  • 适合关注AI公平性与文化安全的研究者

过去三年,大语言模型(LLMs)规模与能力迅速增长,但其训练数据仍以英语为主。多语言大模型的发展日益受到关注,旨在支持全球北方以外社区的语言,包括许多历史上在数字领域被边缘化的语言。这些语言被称为“低资源语言”或“长尾语言”,其在大模型中的表现普遍较差。尽管将大模型应用于更多语言可能促进跨社区沟通与语言保护,但必须警惕数据采集过程中的掠夺性行为,避免重演历史上的剥削实践。对曾被殖民群体、原住民及非西方语言的数据收集,涉及同意、文化安全与数据主权等复杂社会政治问题。此外,语言的复杂性与文化细微差别常在模型中丢失。本文基于最新学术研究与自身工作,梳理了相关社会、文化与伦理考量,并提出通过定性研究、社区合作与参与式设计来缓解这些问题。我们提供了12条关于在非全球北方代表性不足语言社区收集语言数据时应考虑的建议。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have rapidly increased in size and apparent capabilities in the last three years, but their training data is largely English text. There is growing interest in multilingual LLMs, and various efforts are striving for models to accommodate languages of communities outside of the Global North, which include many languages that have been historically underrepresented in digital realms. These languages have been coined as "low resource languages" or "long-tail languages", and LLMs performance on these languages is generally poor. While expanding the use of LLMs to more languages may bring many potential benefits, such as assisting cross-community communication and language preservation, great care must be taken to ensure that data collection on these languages is not extractive and that it does not reproduce exploitative practices of the past. Collecting data from languages spoken by previously colonized people, indigenous people, and non-Western languages raises many complex sociopolitical and ethical questions, e.g., around consent, cultural safety, and data sovereignty. Furthermore, linguistic complexity and cultural nuances are often lost in LLMs. This position paper builds on recent scholarship, and our own work, and outlines several relevant social, cultural, and ethical considerations and potential ways to mitigate them through qualitative research, community partnerships, and participatory design approaches. We provide twelve recommendations for consideration when collecting language data on underrepresented language communities outside of the Global North.

多语言模型伦理数据文化安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。