构建首个马拉地语同义句检测数据集,助力低资源语言理解
MahaParaphrase: A Marathi Paraphrase Detection Corpus and BERT-based Models
- 基于人工标注构建8000对马拉地语句子,区分同义与非同义句
- 在该数据集上,BERT模型达到92.3%的准确率,验证了模型有效性
- 适合研究低资源印地语族语言、自然语言理解及数据增强的学者
同义句是辅助问答、风格迁移、语义解析和数据增强等语言理解任务的重要工具。由于丰富的形态变化、语法差异、多样书写系统以及标注数据稀缺,印地语族语言在自然语言处理中面临挑战。本文提出L3Cube-MahaParaphrase数据集,一个高质量的马拉地语同义句检测语料库,包含8,000对由人工专家标注为同义(P)或非同义(NP)的句子对。同时报告了标准Transformer-based BERT模型在该数据集上的实验结果。数据集与模型已公开于https://github.com/l3cube-pune/MarathiNLP。
原文摘要 · Abstract (English)
Paraphrases are a vital tool to assist language understanding tasks such as question answering, style transfer, semantic parsing, and data augmentation tasks. Indic languages are complex in natural language processing (NLP) due to their rich morphological and syntactic variations, diverse scripts, and limited availability of annotated data. In this work, we present the L3Cube-MahaParaphrase Dataset, a high-quality paraphrase corpus for Marathi, a low resource Indic language, consisting of 8,000 sentence pairs, each annotated by human experts as either Paraphrase (P) or Non-paraphrase (NP). We also present the results of standard transformer-based BERT models on these datasets. The dataset and model are publicly shared at https://github.com/l3cube-pune/MarathiNLP
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。