区分罗曼什语方言的智能识别系统,准确率达97%。
Robust Language Identification for Romansh Varieties
- 用SVM方法构建方言识别模型,支持多种地域变体。
- 在新数据集上实现平均97%的域内识别准确率。
- 适合需要方言感知的拼写检查与机器翻译场景。
罗曼什语有多个地区性变体(称作idioms),部分变体间互懂度较低。尽管存在语言多样性,却缺乏针对这些变体的语言识别(LID)系统研究。由于罗曼什语识别还需识别融合多变体元素的超区域变体Rumantsch Grischun,这构成了一个新颖且有趣的分类问题。本文提出一种基于SVM的罗曼什语方言识别系统,在新构建的基准数据集上跨两个领域进行评估,结果表明其平均域内准确率达到97%,可应用于方言感知的拼写检查或机器翻译。该分类器已公开可用。
原文摘要 · Abstract (English)
The Romansh language has several regional varieties, called idioms, which sometimes have limited mutual intelligibility. Despite this linguistic diversity, there has been a lack of documented efforts to build a language identification (LID) system that can distinguish between these idioms. Since Romansh LID should also be able to recognize Rumantsch Grischun, a supra-regional variety that combines elements of several idioms, this makes for a novel and interesting classification problem. In this paper, we present a LID system for Romansh idioms based on an SVM approach. We evaluate our model on a newly curated benchmark across two domains and find that it reaches an average in-domain accuracy of 97%, enabling applications such as idiom-aware spell checking or machine translation. Our classifier is publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。