为少数语言构建可复用的NLP工具链,推动包容性语言技术发展
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
- 提供从数据采集到下游应用的全流程低资源NLP方法
- 覆盖10+种语言家族与不同地理语境的实证案例
- 强调公平性、可复现性与社区参与,适合实践者与研究者
本教程面向从事多语言与低资源语言NLP的研究人员、开发者与实践者,旨在构建更具公平性与社会影响力的语言技术。参与者将获得一套完整的端到端NLP工具链,涵盖数据收集、网络爬取、平行句对挖掘、机器翻译及文本分类、多模态推理等下游任务。教程针对数据稀缺与文化差异挑战,提供可落地的方法与建模框架,聚焦公平、可复现且以社区为导向的开发范式,基于真实场景展示超过10种来自不同语言家族与地缘背景的语言案例,涵盖数字资源丰富与严重欠代表语言。
原文摘要 · Abstract (English)
This tutorial (https://tum-nlp.github.io/low-resource-tutorial) is designed for NLP practitioners, researchers, and developers working with multilingual and low-resource languages who seek to create more equitable and socially impactful language technologies. Participants will walk away with a practical toolkit for building end-to-end NLP pipelines for underrepresented languages -- from data collection and web crawling to parallel sentence mining, machine translation, and downstream applications such as text classification and multimodal reasoning. The tutorial presents strategies for tackling the challenges of data scarcity and cultural variance, offering hands-on methods and modeling frameworks. We will focus on fair, reproducible, and community-informed development approaches, grounded in real-world scenarios. We will showcase a diverse set of use cases covering over 10 languages from different language families and geopolitical contexts, including both digitally resource-rich and severely underrepresented languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。