arXiv:2601.23022cs.CL2026-01ACL被引 1

构建首个多语言多领域维度情感分析数据集,支持细粒度情感量化。

DimABSA: Building Multilingual and Multidomain Datasets for Dimensional Aspect-Based Sentiment Analysis

  • 用连续情绪值(VA)替代传统分类标签,实现更精细的情感分析。
  • 涵盖6语言4领域,共76,958个情感实例,支持三种新子任务。
  • 提出新评估指标cF1,适合大模型在混合输出任务中测试。

方面级情感分析(ABSA)聚焦于细粒度方面层面的情感提取,已广泛应用于多个现实场景。然而,现有研究依赖粗粒度类别标签(如正面、负面),难以捕捉细微情感状态。为此,我们采用维度化方法,以连续的效价-唤醒度(VA)分数表示情感,从而在方面和情感层面实现更精细分析。为此,我们提出了DimABSA,首个多语言、维度化的ABSA资源,标注了传统ABSA元素(方面词、方面类别、观点词)及新增的VA分数。该资源包含42,590句中的76,958个方面实例,覆盖六种语言和四个领域。我们进一步设计三个子任务,将VA分数与不同ABSA元素结合,搭建从传统到维度化ABSA的桥梁。鉴于这些任务涉及类别与连续输出,我们提出统一指标连续F1(cF1),将VA预测误差纳入标准F1计算。我们在所有子任务上使用提示和微调的大语言模型进行了全面基准测试。结果表明,DimABSA是一个具有挑战性的基准,为推进多语言维度化ABSA奠定了基础。我们已公开发布该数据集,用于SemEval-2026 Task 3 Track A,吸引了超过300名参与者。

原文摘要 · Abstract (English)

Aspect-Based Sentiment Analysis (ABSA) focuses on extracting sentiment at a fine-grained aspect level and has been widely applied across real-world domains. However, existing ABSA research relies on coarse-grained categorical labels (e.g., positive, negative), which limits its ability to capture nuanced affective states. To address this limitation, we adopt a dimensional approach that represents sentiment with continuous valence-arousal (VA) scores, enabling fine-grained analysis at both the aspect and sentiment levels. To this end, we introduce DimABSA, the first multilingual, dimensional ABSA resource annotated with both traditional ABSA elements (aspect terms, aspect categories, and opinion terms) and newly introduced VA scores. This resource contains 76,958 aspect instances across 42,590 sentences, spanning six languages and four domains. We further introduce three subtasks that combine VA scores with different ABSA elements, providing a bridge from traditional ABSA to dimensional ABSA. Given that these subtasks involve both categorical and continuous outputs, we propose a new unified metric, continuous F1 (cF1), which incorporates VA prediction error into standard F1. We provide a comprehensive benchmark using both prompted and fine-tuned large language models across all subtasks. Our results show that DimABSA is a challenging benchmark and provides a foundation for advancing multilingual dimensional ABSA. We publicly released the DimABSA dataset, which was used for Track A of SemEval-2026 Task 3, attracting over 300 participants.

情感分析多语言维度化数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。