用五个相同模型的集成,比复杂架构更有效提升古吉拉特语命名实体识别性能。
HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition?
- 五个微调后的古吉拉特语BERT模型通过投票集成,保持架构统一。
- 集成模型在测试集上达到0.8442的实体级F1,优于单模型和多种异构组合。
- 对低资源语言而言,语言对齐比架构多样性更能提升效果,适合同类研究者参考。
古吉拉特语命名实体识别(NER)因缺乏大写提示、丰富形态、词汇歧义和自由词序而研究不足。以往集成方法强调架构多样性,如混合异构分类器、多语言编码器或传统序列模型,而非利用语言对齐的单语预训练。本研究探讨:对于低资源、形态丰富的古吉拉特语,单一单语编码器的同质集成是否优于架构多样性?提出HomoEnsNER,由五个独立微调的GujaratiBERT模型通过多数投票构成,与单个GujaratiBERT基线及六种异构变体(含MuRIL-base、MuRIL-large、IndicBERT、mBERT、BiLSTM、CRF及堆叠式BiLSTM-CRF-GujaratiBERT)对比。所有模型在相同预算下训练,使用Naamapadam古吉拉特语测试集的实体级F1评估。HomoEnsNER取得最高F1(0.8442),超越基线(0.8347)及所有异构方案(最低:0.7855),表明语言对齐是低资源印度语言NER中更高效、更省预算的集成策略。
原文摘要 · Abstract (English)
Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, and free word order. Prior ensemble work has emphasized architectural diversity by combining heterogeneous classifiers, multilingual encoders, or classical sequence models, rather than exploiting language-aligned monolingual pretraining. This study asks whether, for a low-resource, morphologically rich language like Gujarati, a homogeneous ensemble of a single monolingual encoder outperforms such architectural diversity. We propose HomoEnsNER, a homogeneous ensemble of five independently fine-tuned GujaratiBERT models combined via majority voting, evaluated against a single GujaratiBERT baseline and six heterogeneous alternatives, including combinations with MuRIL-base, MuRIL-large, IndicBERT, mBERT, BiLSTM, CRF, and a stacked BiLSTM-CRF-GujaratiBERT architecture. All eight models were trained under a consistent budget and evaluated using entity-level F1 on the Naamapadam Gujarati test split. HomoEnsNER achieved the highest F1 (0.8442), surpassing the baseline (0.8347) and every heterogeneous alternative (lowest: 0.7855), indicating that language alignment is a more effective, budget-conscious ensembling strategy than architectural complexity for low-resource Indian language NER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。