arXiv:2603.05272cs.CLcs.HC2026-03

为孟加拉国42种濒危语言建立首个全国性多模态语料库。

Oral to Web: Digitizing 'Zero Resource'Languages of Bangladesh

  • 通过90天实地调研采集42种语言的文本与音频数据
  • 涵盖14种濒危语言,含8.5万条结构化条目和107小时录音
  • 适合语言保护、低资源NLP研究者使用

我们构建了首个覆盖孟加拉国民族与原住民语言的全国性多语种平行多模态语料库。尽管该国拥有约40种少数民族语言,分属四个语系,但这些以口语为主、计算上属于‘零资源’的语言长期缺乏系统性数字语料库,其中14种被列为濒危语言。本语料库包含85,792条结构化文本条目,每条包含孟加拉语刺激文本、英文翻译及国际音标(IPA)转写,并附约107小时已转录的音频,涵盖来自藏缅、印欧、南亚、德拉维达语系及两种未分类语言的42种语言变体。数据通过为期90天的实地工作在9个地区收集,涉及16名数据采集员、77位母语者和43名校验员,依据2224个独特项目、三个语言粒度层级的预设提取模板完成:孤立词汇(22个语义领域共475词)、语法结构(21类共887句,含动词变位模式)、定向对话(46个情景共862个提示)。后期处理由10名语言学家进行IPA转写,6名评审独立校对。完整数据集可通过Multilingual Cloud平台(multiling.cloud)公开获取,支持对所有记录语言的标注音频与文本的检索访问。本文详细描述语料库设计、田野方法、数据结构及各语言覆盖率,并探讨其在濒危语言记录、低资源自然语言处理与语言多样性发展中国家数字保存中的意义。

原文摘要 · Abstract (English)

We present the Multilingual Cloud Corpus, the first national-scale, parallel, multimodal linguistic dataset of Bangladesh's ethnic and indigenous languages. Despite being home to approximately 40 minority languages spanning four language families, Bangladesh has lacked a systematic, cross-family digital corpus for these predominantly oral, computationally "zero resource" varieties, 14 of which are classified as endangered. Our corpus comprises 85792 structured textual entries, each containing a Bengali stimulus text, an English translation, and an IPA transcription, together with approximately 107 hours of transcribed audio recordings, covering 42 language varieties from the Tibeto-Burman, Indo-European, Austro-Asiatic, and Dravidian families, plus two genetically unclassified languages. The data were collected through systematic fieldwork over 90 days across nine districts of Bangladesh, involving 16 data collectors, 77 speakers, and 43 validators, following a predefined elicitation template of 2224 unique items organized at three levels of linguistic granularity: isolated lexical items (475 words across 22 semantic domains), grammatical constructions (887 sentences across 21 categories including verbal conjugation paradigms), and directed speech (862 prompts across 46 conversational scenarios). Post-field processing included IPA transcription by 10 linguists with independent adjudication by 6 reviewers. The complete dataset is publicly accessible through the Multilingual Cloud platform (multiling.cloud), providing searchable access to annotated audio and textual data for all documented varieties. We describe the corpus design, fieldwork methodology, dataset structure, and per-language coverage, and discuss implications for endangered language documentation, low-resource NLP, and digital preservation in linguistically diverse developing countries.

语言保护多模态语料低资源NLP濒危语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。