arXiv:2510.20002cs.CLcs.AI2025-10被引 1

针对希腊语数据少、模型旧的问题,打造高质量语料库与新模型。

Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation

  • 用精细清洗和重复优质法律文本提升数据质量
  • 新模型在三个评测中最高提效3.6%,显著优于现有方法
  • 适合研究小语种NLP或法律文本处理的人参考

现代希腊语等形态丰富但资源有限的语言,其自然语言处理进展受限于架构停滞、数据稀缺及领域上下文理解能力不足,尤其在法律等专业领域。本文提出希腊语嵌入模型(GEMs),一套基于Transformer的新型语言模型家族,通过架构多样性和增强数据清洗来应对上述挑战。所用语料库涵盖大规模通用数据集与专用法律语料,采用复杂预处理流程、高级去重策略及高质量法律子语料重复机制,以提升领域适应性。GEMs包含成熟架构(RoBERTa、Longformer)与首次应用于希腊语的先进模型(ELECTRA、ConvBERT、ModernBERT),实现现代Transformer设计的全面覆盖。此外,首次推出面向跨语言法律应用的双语希腊语-英语嵌入模型。在三大核心自然语言理解基准上的综合评估表明,GEM-RoBERTa与GEM-ConvBERT相较已有最先进模型取得统计显著提升,准确率最高提高3.6%;通过Friedman对齐秩检验与Finner事后分析,证实本方法在多指标上均具优势。

原文摘要 · Abstract (English)

The advancement of natural language processing for morphologically rich and moderately-resourced languages like Modern Greek has been hindered by architectural stagnation, data scarcity, and limited context processing capabilities, particularly in specialized domains such as law. In this work, we propose the Greek Embedding Models (GEMs), a new family of transformer-based language models, specifically developed to address these limitations through architectural diversity and enhanced data curation. The proposed family of models are trained on several large-scale, meticulously curated corpora, encompassing both comprehensive general-domain datasets and specialized legal collections, addressing the persistent data scarcity that has impeded Greek language modeling advancement. The proposed quality-based corpus curation methodology incorporates extensive preprocessing pipelines, sophisticated deduplication strategies and targeted repetition of high-quality legal sub-corpora to enhance domain adaptation. The GEMs family comprises both established architectures (RoBERTa and Longformer) and advanced models not previously applied to Greek (ELECTRA, ConvBERT, and ModernBERT), providing comprehensive coverage of modern transformer designs. Additionally, we introduce the first bilingual Greek-English embedding models tailored for cross-lingual legal applications. Comprehensive evaluation across three core natural language understanding benchmarks demonstrates that the proposed GEM-RoBERTa and GEM-ConvBERT achieve statistically significant performance improvements over established state-of-the-art models, with accuracy gains of up to 3.6\% while conducted statistical analysis using Friedman Aligned-Ranks and Finner post-hoc tests confirms the superiority of our approach across multiple evaluation metrics.

希腊语法律NLP数据清洗Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。