为塞尔维亚语等低资源语言构建文化敏感的语言技术框架
From Data Scarcity to Data Care: Reimagining Language Technologies for Serbian and other Low-Resource Languages
- 提出Data Care框架,将伦理与集体控制融入语料设计
- 揭示历史破坏与工程优先导致的语料偏差问题
- 适合关注文化公平与语言多样性研究者参考
大型语言模型通常基于英语等主流语言训练,其对低资源语言的表征常反映源语言材料中的文化与语言偏见。以塞尔维亚语为例,本研究通过与十位学者及从业者(包括语言学家、数字人文学者和AI开发者)的半结构化访谈,探讨了影响低资源语言技术发展的结构性、历史性和社会技术因素。研究指出,历史上的文本遗产破坏与当前工程导向的简化方法共同导致了浅层音译、依赖英语模型、数据偏见及缺乏文化特异性的数据集构建。为此,本文提出基于CARE原则(集体利益、控制权、责任与伦理)的Data Care框架,将偏见缓解从事后技术修补转变为语料设计、标注与治理的核心环节,为在传统大模型开发加剧权力失衡与文化盲点的背景下,构建包容、可持续且文化根基深厚的语言技术提供可复制范式。
原文摘要 · Abstract (English)
Large language models are commonly trained on dominant languages like English, and their representation of low resource languages typically reflects cultural and linguistic biases present in the source language materials. Using the Serbian language as a case, this study examines the structural, historical, and sociotechnical factors shaping language technology development for low resource languages in the AI age. Drawing on semi structured interviews with ten scholars and practitioners, including linguists, digital humanists, and AI developers, it traces challenges rooted in historical destruction of Serbian textual heritage, intensified by contemporary issues that drive reductive, engineering first approaches prioritizing functionality over linguistic nuance. These include superficial transliteration, reliance on English-trained models, data bias, and dataset curation lacking cultural specificity. To address these challenges, the study proposes Data Care, a framework grounded in CARE principles (Collective Benefit, Authority to Control, Responsibility, and Ethics), that reframes bias mitigation from a post hoc technical fix to an integral component of corpus design, annotation, and governance, and positions Data Care as a replicable model for building inclusive, sustainable, and culturally grounded language technologies in contexts where traditional LLM development reproduces existing power imbalances and cultural blind spots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。