梳理南亚21种低资源语言的文本、语音与多模态研究进展
A Breadth-First Catalog of Text Processing, Speech Processing and Multimodal Research in South Asian Languages
- 基于大模型对南亚语言研究进行分类与聚类分析
- 覆盖2022年1月至2024年10月期间的最新文献
- 聚焦21种低资源语言,为研究者提供全景视角
本文综述了2022年1月至2024年10月间南亚语言在文本处理、多模态模型和语音处理领域的最新研究,重点关注21种低资源语言:Saraiki、Assamese、Balochi、Bhojpuri、Bodo、Burmese、Chhattisgarhi、Dhivehi、Gujarati、Kannada、Kashmiri、Konkani、Khasi、Malayalam、Meitei、Nepali、Odia、Pashto、Rajasthani、Sindhi和Telugu。采用基于大语言模型(LLMs)的相关性分类与聚类方法,识别研究趋势、挑战与未来方向。旨在为关注南亚语言技术的NLP研究者提供一份广度优先的近期发展概览。
原文摘要 · Abstract (English)
We review the recent literature (January 2022- October 2024) in South Asian languages on text-based language processing, multimodal models, and speech processing, and provide a spotlight analysis focused on 21 low-resource South Asian languages, namely Saraiki, Assamese, Balochi, Bhojpuri, Bodo, Burmese, Chhattisgarhi, Dhivehi, Gujarati, Kannada, Kashmiri, Konkani, Khasi, Malayalam, Meitei, Nepali, Odia, Pashto, Rajasthani, Sindhi, and Telugu. We identify trends, challenges, and future research directions, using a step-wise approach that incorporates relevance classification and clustering based on large language models (LLMs). Our goal is to provide a breadth-first overview of the recent developments in South Asian language technologies to NLP researchers interested in working with South Asian languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。