构建多语言语义框架库,提升跨语言语义解析性能。
The Multilingual FrameNet Corpus
- 整合九种语言的语料,统一构建多语言语义框架库。
- 在多语言与跨语言任务中均超越现有最优模型。
- 适合从事多语言语义分析与跨语言自然语言处理的研究者。
本文介绍多语言语义框架语料库(mFNC),该资源在英语伯克利语义框架语料库基础上,收集并统一了巴西葡萄牙语、中文、荷兰语、法语、德语、意大利语、韩语、拉脱维亚语和瑞典语共九种语言的已有语料。在mFNC上训练不同架构的语义解析模型,均在多语言及跨语言设置下持续优于现有最先进方法,凸显多语言训练数据的重要性。mFNC及训练好的语义解析模型已开源,可访问 https://github.com/beatrice-f/mFNC。
原文摘要 · Abstract (English)
This paper introduces the Multilingual FrameNet Corpus (mFNC), a novel resource that extends the English Berkeley FrameNet corpus by collecting and harmonizing existing language-specific corpora across nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian and Swedish. By training models that rely on different architectures on the mFNC, we consistently outperform existing state-of-the-art Frame Semantic Parsers in both multilingual and cross-lingual settings, underscoring the importance of multilingual training data. The mFNC and our trained FSP models are openly available at https://github.com/beatrice-f/mFNC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。