构建多方言阿拉伯语情感分析数据集,助力酒店业客户反馈分析
AHaSIS: Shared Task on Sentiment Analysis for Arabic Dialects
- 基于阿语方言酒店评论构建538条平衡数据集
- 最佳模型F1达0.81,验证跨方言分析可行性
- 适合关注中东NLP与客户服务分析的研究者
阿拉伯世界旅游业日益依赖客户反馈优化服务,亟需先进的阿拉伯语情感分析工具。为此,本次共享任务聚焦阿拉伯语方言在酒店领域的意见检测。任务基于多方言、人工标注的数据集,该数据集源自原始现代标准阿拉伯语(MSA)的酒店评论,并翻译为沙特阿拉伯语和摩洛哥方言(Darija)。数据集包含538条情感均衡的评论,涵盖正面、中性、负面三类。翻译经母语者验证,确保方言准确性和情感一致性。该资源支持开发面向真实场景的方言感知自然语言处理系统。超过40支队伍注册,12支提交系统。最优模型取得0.81的F1分数,证明跨方言情感分析的可行性与现存挑战。
原文摘要 · Abstract (English)
The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address this challenge, the Sentiment Analysis on Arabic Dialects in the Hospitality Domain shared task focuses on Sentiment Detection in Arabic Dialects. This task leverages a multi-dialect, manually curated dataset derived from hotel reviews originally written in Modern Standard Arabic (MSA) and translated into Saudi and Moroccan (Darija) dialects. The dataset consists of 538 sentiment-balanced reviews spanning positive, neutral, and negative categories. Translations were validated by native speakers to ensure dialectal accuracy and sentiment preservation. This resource supports the development of dialect-aware NLP systems for real-world applications in customer experience analysis. More than 40 teams have registered for the shared task, with 12 submitting systems during the evaluation phase. The top-performing system achieved an F1 score of 0.81, demonstrating the feasibility and ongoing challenges of sentiment analysis across Arabic dialects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。