arXiv:2601.09648cs.CL2026-01中稿 · LREC 2026被引 1

用银标准数据构建多语言语义标注系统,融合规则与神经网络提升性能。

Creating a Hybrid Rule and Neural Network Based Semantic Tagger using Silver Standard Data: the PyMUSAS framework for Multilingual Semantic Annotation

  • 基于规则系统结合神经网络模型,利用银标准数据训练新模型。
  • 在五种语言上评估,中文新数据集提升跨语言标注效果。
  • 开源代码、数据与模型,支持多语言语义分析研究。

词义消歧(WSD)通常基于WordNet、BabelNet和牛津英语词典等语义框架进行评估。然而,对于UCREL语义分析系统(USAS)框架,尚未开展大规模公开评估,也缺乏多语言覆盖的全面验证。本文首次对基于规则的USAS系统进行了最大规模的语义标注评估,覆盖五种语言,使用四个现有数据集及一个全新的中文数据集。为解决人工标注数据不足问题,我们构建了一个新的银标准英文数据集,训练并评估了多种单语与多语神经网络模型,在单语与跨语言设置下与规则系统对比,证明神经网络可有效增强规则系统。所提出的神经模型、训练数据、中文评估集及全部代码均已开源。

原文摘要 · Abstract (English)

Word Sense Disambiguation (WSD) has been widely evaluated using the semantic frameworks of WordNet, BabelNet, and the Oxford Dictionary of English. However, for the UCREL Semantic Analysis System (USAS) framework, no open extensive evaluation has been performed beyond lexical coverage or single language evaluation. In this work, we perform the largest semantic tagging evaluation of the rule based system that uses the lexical resources in the USAS framework covering five different languages using four existing datasets and one novel Chinese dataset. We create a new silver labelled English dataset, to overcome the lack of manually tagged training data, that we train and evaluate various mono and multilingual neural models in both mono and cross-lingual evaluation setups with comparisons to their rule based counterparts, and show how a rule based system can be enhanced with a neural network model. The resulting neural network models, including the data they were trained on, the Chinese evaluation dataset, and all of the code have been released as open resources.

语义标注多语言神经网络开源数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。