首个基于分类体系的斯瓦希里语自然语言评估,揭示方言多样性对模型表现的影响。
Benchmarking Sociolinguistic Diversity in Swahili NLP: A Taxonomy-Guided Approach
- 构建分类体系,系统分析斯瓦希里语中的部落影响、城市口语等语言变异
- 在2170条肯尼亚人自由文本数据上发现模型在混合语码和借词上的错误率显著上升
- 为非洲语言NLP提供文化敏感的评估框架,适合关注语言公平性的研究者
我们提出首个基于分类体系的斯瓦希里语自然语言处理评估方法,弥补社会语言多样性方面的空白。基于与健康相关的心理测量任务,收集了来自肯尼亚说话者的2,170条自由文本回答,数据中包含部落影响、城市俚语、语码混用和外来词。我们构建了结构化分类体系,并以此作为分析预训练和指令微调语言模型预测错误的视角。研究结果推动了文化根植的评估框架发展,凸显社会语言变体对模型性能的关键影响。
原文摘要 · Abstract (English)
We introduce the first taxonomy-guided evaluation of Swahili NLP, addressing gaps in sociolinguistic diversity. Drawing on health-related psychometric tasks, we collect a dataset of 2,170 free-text responses from Kenyan speakers. The data exhibits tribal influences, urban vernacular, code-mixing, and loanwords. We develop a structured taxonomy and use it as a lens for examining model prediction errors across pre-trained and instruction-tuned language models. Our findings advance culturally grounded evaluation frameworks and highlight the role of sociolinguistic variation in shaping model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。