arXiv:2503.10298cs.CL2025-03

探讨大模型语言多样性缺失及其对技术公平性的影响

Proceedings of the ISCA/ITG Workshop on Diversity in Large Speech and Language Models

  • 从计算机科学与语言学视角分析大模型训练数据的语种偏见
  • 指出低资源语言因数据匮乏面临技术边缘化风险
  • 适合关注AI公平性、语言多样性研究的读者

机器学习已广泛应用于语音与自然语言处理任务,如语音识别、信息抽取、文本与语音生成及人机交互(如聊天机器人)。现代技术依赖大规模模型来表征一种或多种语言的通用知识(大型语言模型,LLMs)或语音与音频特征。这些模型通常使用大量网络内容进行训练。当人类与这类技术互动时,交互效果取决于用户使用的语言是否与模型训练所用语言一致;若模型无法泛化至用户实际语言,则可能导致用户逐渐适应模型语言,不适应者将被排除在高效使用之外。此外,商业模型开发受市场需求驱动,导致低代表语言和方言/社会方言的优先级下降。对于许多使用人数较少的语言,缺乏必要数据进一步加剧了语音与语言技术应用中的数字鸿沟。本次研讨会基于计算机科学与语言学(包括计算语言学与NLP)的科学贡献,讨论该问题。

原文摘要 · Abstract (English)

Machine learning techniques have conquered many different tasks in speech and natural language processing, such as speech recognition, information extraction, text and speech generation, and human machine interaction using natural language or speech (chatbots). Modern techniques typically rely on large models for representing general knowledge of one or several languages (Large Language Models, LLMs), or for representing speech and general audio characteristics. These models have been trained with large amounts of speech and language data, typically including web content. When humans interact with such technologies, the effectiveness of the interaction will be influenced by how far humans make use of the same type of language the models have been trained on or, in other words, if the models are able to generalize to the language used by humans when interacting with the technology. This may lead to some gradual forms of adaptation in human speech and language production, and users who do not adapt may be excluded from efficient use of such technologies. On top of this, as commercial model development follows market needs, under-represented languages and dialects/sociolects may decrease in terms of priorities. Furthermore, for many lesser spoken languages the necessary data is not available, which will worsen a digital divide in speech and language technology usage. The workshop sets out to discuss this problem based on scientific contributions from the perspective of computer science and linguistics (including computational linguistics and NLP).

语言多样性大模型数字鸿沟AI公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。