arXiv:2509.16788cs.CLcs.AI2025-09被引 4

针对阿拉伯语情感分析,提出领域自适应预训练新方法。

Domain-Adaptive Pre-Training for Arabic Aspect-Based Sentiment Analysis: A Comparative Study of Domain Adaptation and Fine-Tuning Strategies

  • 用领域自适应预训练提升阿拉伯语情感分析模型
  • 适配器微调在效率与性能间表现最佳
  • 发现标注与模型预测存在系统性偏差

基于方面的情感分析(ABSA)可帮助机构理解用户对产品特定方面的评价。尽管深度学习模型广泛应用于英文ABSA,但阿拉伯语因标注数据稀缺而应用受限。现有研究尝试使用BERT等上下文感知预训练模型,但这些模型多基于事实性数据,可能引入领域偏见。目前尚无研究将自适应预训练应用于阿拉伯语上下文模型的ABSA任务。本研究提出一种新的领域自适应预训练方法,用于方面情感分类(ASC)和观点目标抽取(OTE)。比较了特征提取、全量微调与适配器微调三种策略,在多个适配语料库与上下文模型上进行实验。结果表明,领域内自适应预训练带来适度提升;适配器微调在计算效率与性能上表现优异。但误差分析揭示模型预测与数据标注存在诸多问题:在ASC中,常见错误包括情感标签误标、对比标记误读、早期词项正向偏倚,以及矛盾意见和子词切分挑战;在OTE中,则表现为目标误标、句法角色混淆、多词表达识别困难,以及依赖浅层启发式规则。这些发现强调需发展具备语法与语义感知能力的模型,如图卷积网络,以更有效捕捉长距离关系与复杂方面-观点对齐。

原文摘要 · Abstract (English)

Aspect-based sentiment analysis (ABSA) in natural language processing enables organizations to understand customer opinions on specific product aspects. While deep learning models are widely used for English ABSA, their application in Arabic is limited due to the scarcity of labeled data. Researchers have attempted to tackle this issue by using pre-trained contextualized language models such as BERT. However, these models are often based on fact-based data, which can introduce bias in domain-specific tasks like ABSA. To our knowledge, no studies have applied adaptive pre-training with Arabic contextualized models for ABSA. This research proposes a novel approach using domain-adaptive pre-training for aspect-sentiment classification (ASC) and opinion target expression (OTE) extraction. We examine fine-tuning strategies - feature extraction, full fine-tuning, and adapter-based methods - to enhance performance and efficiency, utilizing multiple adaptation corpora and contextualized models. Our results show that in-domain adaptive pre-training yields modest improvements. Adapter-based fine-tuning is a computationally efficient method that achieves competitive results. However, error analyses reveal issues with model predictions and dataset labeling. In ASC, common problems include incorrect sentiment labeling, misinterpretation of contrastive markers, positivity bias for early terms, and challenges with conflicting opinions and subword tokenization. For OTE, issues involve mislabeling targets, confusion over syntactic roles, difficulty with multi-word expressions, and reliance on shallow heuristics. These findings underscore the need for syntax- and semantics-aware models, such as graph convolutional networks, to more effectively capture long-distance relations and complex aspect-based opinion alignments.

阿拉伯语情感分析自适应预训练适配器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。