用AI和语言学结合,抢救印度濒危的托托语
Integrating Linguistics and AI: Morphological Analysis and Corpus development of Endangered Toto Language of West Bengal
- 通过田野调查构建带词素标注的三语语料库
- 训练小型语言模型与翻译引擎,支持语言学习应用
- 适合语言保护者与跨学科研究者参考
保护语言多样性至关重要,每种语言都提供了独特的世界观。本文是旨在开发三语(托托-孟加拉-英语)语言学习应用的项目一部分,以数字化存档并推广印度西孟加拉邦的濒危语言托托语。该应用面向母语者与非母语学习者,通过集成Unicode字符集和结构化语料库提升可访问性与可用性。研究包含实地采集的语言学资料,建立词素标注的三语语料库,用于训练小型语言模型(SLM)与基于Transformer的翻译引擎。分析涵盖屈折形态如人称-数-性一致、时态-体-语气区分、格标记,以及体现词类变化的派生策略。同时完成书写系统标准化与数字识字工具开发。本研究提出可持续的语言保护模式,融合传统语言学方法与人工智能技术,彰显跨学科合作在社区主导语言复兴中的价值。
原文摘要 · Abstract (English)
Preserving linguistic diversity is necessary as every language offers a distinct perspective on the world. There have been numerous global initiatives to preserve endangered languages through documentation. This paper is a part of a project which aims to develop a trilingual (Toto-Bangla-English) language learning application to digitally archive and promote the endangered Toto language of West Bengal, India. This application, designed for both native Toto speakers and non-native learners, aims to revitalize the language by ensuring accessibility and usability through Unicode script integration and a structured language corpus. The research includes detailed linguistic documentation collected via fieldwork, followed by the creation of a morpheme-tagged, trilingual corpus used to train a Small Language Model (SLM) and a Transformer-based translation engine. The analysis covers inflectional morphology such as person-number-gender agreement, tense-aspect-mood distinctions, and case marking, alongside derivational strategies that reflect word-class changes. Script standardization and digital literacy tools were also developed to enhance script usage. The study offers a sustainable model for preserving endangered languages by incorporating traditional linguistic methodology with AI. This bridge between linguistic research with technological innovation highlights the value of interdisciplinary collaboration for community-based language revitalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。