arXiv:2601.00411cs.CLcs.AI2026-01中稿 · LREC 2026被引 1

用大模型验证弱监督标注,构建卢森堡语命名实体识别新数据集

Do LLMs Judge Distantly Supervised Named Entity Labels Well? Constructing the JudgeWEL Dataset

  • 利用维基百科和Wikidata的内部链接生成弱监督标注
  • 通过多LLM对比筛选,使数据量达现有数据集5倍且覆盖更均衡
  • 为低资源语言NER研究提供高质量新基准,适合多语言模型开发者

我们提出judgeWEL,一个用于卢森堡语命名实体识别(NER)的数据集,通过大语言模型(LLM)在新型流水线中自动标注并验证。构建低资源语言数据集仍是自然语言处理的主要瓶颈,因资源稀缺与语言特殊性导致大规模人工标注成本高且易不一致。为此,我们提出并评估一种新方法:利用维基百科和Wikidata作为结构化弱监督源。通过维基百科文章内的内部链接,基于其对应的Wikidata条目推断实体类型,从而实现最小人工干预下的初始标注。由于此类链接可靠性不均,我们采用并比较多种LLM,仅保留高质量标注句子。最终语料规模约为当前可用卢森堡语NER数据集的五倍,且在实体类别上覆盖更广、分布更均衡,为多语言和低资源NER研究提供了重要新资源。

原文摘要 · Abstract (English)

We present judgeWEL, a dataset for named entity recognition (NER) in Luxembourgish, automatically labelled and subsequently verified using large language models (LLM) in a novel pipeline. Building datasets for under-represented languages remains one of the major bottlenecks in natural language processing, where the scarcity of resources and linguistic particularities make large-scale annotation costly and potentially inconsistent. To address these challenges, we propose and evaluate a novel approach that leverages Wikipedia and Wikidata as structured sources of weak supervision. By exploiting internal links within Wikipedia articles, we infer entity types based on their corresponding Wikidata entries, thereby generating initial annotations with minimal human intervention. Because such links are not uniformly reliable, we mitigate noise by employing and comparing several LLMs to identify and retain only high-quality labelled sentences. The resulting corpus is approximately five times larger than the currently available Luxembourgish NER dataset and offers broader and more balanced coverage across entity categories, providing a substantial new resource for multilingual and low-resource NER research.

命名实体识别低资源语言大模型标注多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。