聚焦南半球多语言边缘部署,推动包容性语言模型发展
Multilinguality at the Edge: Developing Language Models for the Global South
- 跨语言与边缘计算融合研究,破解资源匮乏地区技术落地难题
- 调研232篇论文,系统梳理从数据到部署全链条挑战
- 为开发者、政策制定者提供可操作建议,促进技术公平
语言模型的部署位置和方式决定了谁能从中受益。然而,非英语语种且硬件受限的全球南方社区面临诸多挑战,这被称为‘最后一公里’问题:多语言与边缘部署的交汇点,目标一致但技术需求常冲突。我们通过调研232篇涵盖语言建模全流程(从数据收集到开发与部署)的论文,深入分析这一交叉领域的现状与挑战。同时探讨开放问题,并为自然语言处理生态中不同利益相关方提出具体行动建议。期望本工作能推动更具包容性和公平性的语言技术发展。
原文摘要 · Abstract (English)
Where and how language models (LMs) are deployed determines who can benefit from them. However, there are several challenges that prevent effective deployment of LMs in non-English-speaking and hardware constrained communities in the Global South. We call this challenge the last mile: the intersection of multilinguality and edge deployment, where the goals are aligned but the technical requirements often compete. Studying these two fields together is both a need, as linguistically diverse communities often face the most severe infrastructure constraints, and an opportunity, as edge and multilingual NLP research remain largely siloed. To understand the state of the art and the challenges of combining the two areas, we survey 232 papers that tackle this problem across the language modelling pipeline, from data collection to development and deployment. We also discuss open questions and provide actionable recommendations for different stakeholders in the NLP ecosystem. Finally, we hope that this work contributes to the development of inclusive and equitable language technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。