arXiv:2409.06550cs.CL2024-09被引 1

LIMA框架融合深度学习模型,实现60+语言的文本分析互操作。

From LIMA to DeepLIMA: following a new path of interoperability

  • 基于深度神经网络扩展多语言分析模块
  • 在60多个语言上训练模型,支持跨平台集成
  • 通过标准化数据与模型提升系统间兼容性

本文描述了LIMA(Libre Multilingual Analyzer)框架的架构及其最新演进,新增基于深度神经网络的文本分析模块。在保持原有可配置架构和已有规则/统计分析组件的基础上,扩展了对更多语言的支持。模型在超过60种语言上基于Universal Dependencies 2.5、WikiNer和CoNLL-03数据集进行训练。Universal Dependencies的使用不仅增加了支持语言数量,还使模型可被其他平台集成。将通用深度学习自然语言处理模型与标准标注语料结合,构成一种新型互操作路径,通过模型与数据的标准化,补充了原本由Docker容器服务实现的技术互操作性。

原文摘要 · Abstract (English)

In this article, we describe the architecture of the LIMA (Libre Multilingual Analyzer) framework and its recent evolution with the addition of new text analysis modules based on deep neural networks. We extended the functionality of LIMA in terms of the number of supported languages while preserving existing configurable architecture and the availability of previously developed rule-based and statistical analysis components. Models were trained for more than 60 languages on the Universal Dependencies 2.5 corpora, WikiNer corpora, and CoNLL-03 dataset. Universal Dependencies allowed us to increase the number of supported languages and to generate models that could be integrated into other platforms. This integration of ubiquitous Deep Learning Natural Language Processing models and the use of standard annotated collections using Universal Dependencies can be viewed as a new path of interoperability, through the normalization of models and data, that are complementary to a more standard technical interoperability, implemented in LIMA through services available in Docker containers on Docker Hub.

多语言分析深度学习互操作性NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。