arXiv:2604.26230cs.CLstat.ME2026-04

用词向量做掩码模型改进情感分析,更准更易懂。

A New Semisupervised Technique for Polarity Analysis using Masked Language Models

  • 把词向量当掩码语言模型,用种子词出现概率给文本打分
  • 在疫情报道中比传统方法更准确、一致、可解释
  • 适合想提升半监督情感分析效果的研究者

本文提出一种基于word2vec的新型潜在语义缩放(LSS)方法,将词向量作为掩码语言模型使用。与传统空间模型不同,该方法通过预测种子词在特定上下文中的出现概率,为词语和文档赋予极性分数。这些概率型极性评分在准确性、可解释性和一致性上均优于传统空间模型。通过将该方法应用于《中国日报》在新冠疫情期间对中美等国在健康领域成就的报道分析,验证了其优势。结果表明,采用更先进的掩码语言模型将进一步提升半监督机器学习技术的效果。

原文摘要 · Abstract (English)

I developed a new version of Latent Semantic Scaling (LSS) employing word2vec as a masked language model. Unlike original spatial models, it assigns polarity scores to words and documents as predicted probabilities of seed words to occur in given contexts. These probabilistic polarity scores are more accurate, interpretable and consistent than those spatial polarity models can produce in text analysis. I demonstrate these advantages by applying both probabilistic and spatial models to China Daily's coverage of China and other countries during the coronavirus disease (COVID) pandemic in terms of achievement in health issues. The result suggests that more advanced masked language models would further improve the semisupervised machine learning technique.

情感分析半监督掩码模型词向量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。