动态调节领域信息,提升多领域情感分类效果
Dynamic Domain Information Modulation Algorithm for Multi-domain Sentiment Analysis
- 分两阶段训练:先确定共享超参数,再动态调节领域信息
- 在16个领域的公开数据集上准确率显著优于现有方法
- 适合需要跨领域情感分析的工业应用与研究者
多领域情感分类旨在通过利用多个领域的标注数据,缓解单一领域标注数据稀缺导致的模型性能下降问题。现有联合训练领域分类器和情感分类器的方法虽有效,但其依赖可调权重或超参数控制领域信息影响,当领域数量增加时面临计算资源消耗大、收敛困难和算法复杂度高等挑战。为此,本文提出一种动态信息调制算法。模型训练分为两个阶段:第一阶段确定一个共享超参数,用于控制各领域中领域分类任务的占比;第二阶段引入基于梯度和损失的领域感知调制机制,动态调整输入文本中的领域信息。在包含16个领域的公开情感分析数据集上的实验结果表明,所提方法在性能上具有明显优势。
原文摘要 · Abstract (English)
Multi-domain sentiment classification aims to mitigate poor performance models due to the scarcity of labeled data in a single domain, by utilizing data labeled from various domains. A series of models that jointly train domain classifiers and sentiment classifiers have demonstrated their advantages, because domain classification helps generate necessary information for sentiment classification. Intuitively, the importance of sentiment classification tasks is the same in all domains for multi-domain sentiment classification; but domain classification tasks are different because the impact of domain information on sentiment classification varies across different fields; this can be controlled through adjustable weights or hyper parameters. However, as the number of domains increases, existing hyperparameter optimization algorithms may face the following challenges: (1) tremendous demand for computing resources, (2) convergence problems, and (3) high algorithm complexity. To efficiently generate the domain information required for sentiment classification in each domain, we propose a dynamic information modulation algorithm. Specifically, the model training process is divided into two stages. In the first stage, a shared hyperparameter, which would control the proportion of domain classification tasks across all fields, is determined. In the second stage, we introduce a novel domain-aware modulation algorithm to adjust the domain information contained in the input text, which is then calculated based on a gradient-based and loss-based method. In summary, experimental results on a public sentiment analysis dataset containing 16 domains prove the superiority of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。