arXiv:2604.17134cs.CL2026-04AAAI

构建罗语意语多领域情感分析数据集,提升跨语言跨域模型性能

RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

论文配图:RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
图 1 · 摘自论文原文
  • 提出多目标对抗训练框架,动态平衡情感判别与语言域不变性
  • XLM-R在新数据集上F1达66.23%,比基线提升4.64%
  • 适合研究跨语言情感分析的开发者和多语言NLP研究者

我们提出了RoIt-XMASA,一个将跨语言多领域亚马逊情感分析扩展至意大利语和罗马尼亚语的数据集,包含36,000条标注评论,覆盖书籍、电影和音乐三个领域,以及202,141条未标注样本。为应对跨语言和跨域挑战,我们提出一种多目标对抗训练框架,采用元学习系数的损失反转机制,动态平衡情感判别与语言和领域不变性。实验表明,XLM-R在该方法下达到66.23%的F1分数,较基线提升4.64%。少样本评估显示,Llama-3.1-8B在提示工程下获得58.43% F1,揭示了提示方法效率与任务微调性能之间的权衡。

原文摘要 · Abstract (English)

We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising 36,000 labeled reviews across three domains (books, movies, and music) and 202,141 unlabeled samples. To address cross-lingual and cross-domain challenges, we propose a multi-target adversarial training framework that employs loss reversal with meta-learned coefficients to dynamically balance sentiment discrimination with domain and language invariance. XLM-R achieves an F1-score of 66.23% with our approach, outperforming the baseline by 4.64%. Few-shot evaluation shows that Llama-3.1-8B achieves 58.43% F1-score, revealing a meaningful trade-off between the efficiency of prompting-based approaches and the higher performance of task-specific fine-tuning.

情感分析多语言跨域数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。