首个自动识别伦巴第方言拼写变体的方法与数据集
LombardoGraphia: Automatic Classification of Lombard Orthography Variants
- 构建包含11,186条文本的伦巴第语拼写变体标注语料库
- 最佳模型整体准确率达96.06%,少数类仍受数据不均衡影响
- 为资源匮乏的伦巴第语构建多样性感知的NLP工具提供基础
伦巴第语是意大利北部和瑞士南部约380万人口使用的语言,但缺乏统一的拼写标准,存在多种拼写系统,给自然语言处理资源开发和模型训练带来挑战。本文首次开展伦巴第语拼写变体的自动分类研究,提出LombardoGraphia,一个包含11,186条伦巴第语维基百科样本的标注语料库,覆盖9种拼写变体。通过清洗和过滤原始维基内容,确保文本适合拼写分析。我们训练了24种传统与神经网络分类模型,采用不同特征和编码方式。最优模型在整体准确率上达到96.06%,平均类别准确率为85.78%,但少数类表现受限于数据不平衡问题。本工作为构建面向语言变体的伦巴第语NLP资源提供了关键基础设施。
原文摘要 · Abstract (English)
Lombard, an underresourced language variety spoken by approximately 3.8 million people in Northern Italy and Southern Switzerland, lacks a unified orthographic standard. Multiple orthographic systems exist, creating challenges for NLP resource development and model training. This paper presents the first study of automatic Lombard orthography classification and LombardoGraphia, a curated corpus of 11,186 Lombard Wikipedia samples tagged across 9 orthographic variants, and models for automatic orthography classification. We curate the dataset, processing and filtering raw Wikipedia content to ensure text suitable for orthographic analysis. We train 24 traditional and neural classification models with various features and encoding levels. Our best models achieve 96.06% and 85.78% overall and average class accuracy, though performance on minority classes remains challenging due to data imbalance. Our work provides crucial infrastructure for building variety-aware NLP resources for Lombard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。