用少量数据实现阿拉伯语方言酒店评论情感分析
MAPROC at AHaSIS Shared Task: Few-Shot and Sentence Transformer for Sentiment Analysis of Arabic Hotel Reviews
- 采用高效少样本学习框架SetFit微调句向量模型
- 在官方测试集上达到73%的F1分数
- 适合资源稀缺下处理特定领域阿拉伯语文本
阿拉伯语方言的情感分析因语言多样性及标注数据稀缺而面临重大挑战。本文描述了我们在AHaSIS共享任务中的方法,该任务聚焦于酒店领域阿拉伯语方言的情感分析。数据集包含摩洛哥和沙特方言的酒店评论,目标是将评论情感分类为正面、负面或中性。我们采用了SetFit(句向量微调)这一数据高效的少样本学习技术。在官方评估集上,系统F1得分为73%,在26名参赛者中排名第12。本工作凸显了少样本学习在应对专业领域(如酒店评论)中复杂方言文本数据稀缺问题上的潜力。
原文摘要 · Abstract (English)
Sentiment analysis of Arabic dialects presents significant challenges due to linguistic diversity and the scarcity of annotated data. This paper describes our approach to the AHaSIS shared task, which focuses on sentiment analysis on Arabic dialects in the hospitality domain. The dataset comprises hotel reviews written in Moroccan and Saudi dialects, and the objective is to classify the reviewers sentiment as positive, negative, or neutral. We employed the SetFit (Sentence Transformer Fine-tuning) framework, a data-efficient few-shot learning technique. On the official evaluation set, our system achieved an F1 of 73%, ranking 12th among 26 participants. This work highlights the potential of few-shot learning to address data scarcity in processing nuanced dialectal Arabic text within specialized domains like hotel reviews.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。