构建首个乌尔都语方言比喻语言数据集并评估模型表现
SiNFluD: Creating and Evaluating Figurative Language Dataset for Sindhi
- 从博客等多源收集文本,双标注员标注并达0.81一致性
- 使用交叉验证测试模型,XLM-RoBERTa-XL表现最佳
- 适合研究低资源语言、比喻识别与跨语言模型的学者
本文提出SiNFluD,首个针对信德语比喻语言分类的基准数据集。我们从博客、社交媒体和文学作品中收集原始文本,并进行标注准备。两名母语标注员使用Doccano工具标注数据,达成0.81的标注者间一致性。通过5折和10折交叉验证建立基线结果,评估mBERT、XLM-RoBERTa及XLM-RoBERTa-XL模型,以及SetFit在句子变换器上的少样本微调能力。其中,预训练的XLM-RoBERTa-XL表现最优。
原文摘要 · Abstract (English)
In this article, we introduce SiNFluD, a novel benchmark dataset for Sindhi figurative language classification. We first collect raw text from various blogs, social media platforms, and literary sources, and subsequently prepare the corpus for annotation. Two native annotators label the data using the Doccano text annotation tool, achieving an inter-annotator agreement of 0.81. We then establish baseline results using 5-fold and 10-fold cross-validation. Finally, we evaluate mBERT, XLM-RoBERTa, and XLM-RoBERTa-XL models, along with SetFit for few-shot fine-tuning of sentence transformers. Among these, the pretrained XLM-RoBERTa-XL achieves the best performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。