BERT比词向量更稳定,适合分析政治文本语义变化。
Achieving Semantic Consistency: Contextualized Word Representations for Political Text Analysis
- 用BERT的上下文词向量替代静态词嵌入
- 在20年人民日报数据中表现更稳定且能捕捉细微变化
- 适合需语义稳定的文本分析任务
准确理解词汇在政治科学文本分析中至关重要;某些任务假设语义稳定,而另一些则旨在追踪语义演变。传统静态词嵌入(如Word2Vec)虽能有效捕捉长期语义变化,但因训练数据不平衡常导致短期上下文中嵌入波动,缺乏稳定性。BERT采用基于Transformer的架构和上下文嵌入,表现出更强的语义一致性,适用于对稳定性要求高的分析。本研究利用20年《人民日报》文章数据,对比Word2Vec与BERT在不同时段的语义表征性能。结果表明,BERT在保持语义稳定性方面优于Word2Vec,同时仍能识别细微的语义变化。这些发现支持在不假设语义改变的任务中使用BERT,其可靠性高于静态模型。
原文摘要 · Abstract (English)
Accurately interpreting words is vital in political science text analysis; some tasks require assuming semantic stability, while others aim to trace semantic shifts. Traditional static embeddings, like Word2Vec effectively capture long-term semantic changes but often lack stability in short-term contexts due to embedding fluctuations caused by unbalanced training data. BERT, which features transformer-based architecture and contextual embeddings, offers greater semantic consistency, making it suitable for analyses in which stability is crucial. This study compares Word2Vec and BERT using 20 years of People's Daily articles to evaluate their performance in semantic representations across different timeframes. The results indicate that BERT outperforms Word2Vec in maintaining semantic stability and still recognizes subtle semantic variations. These findings support BERT's use in text analysis tasks that require stability, where semantic changes are not assumed, offering a more reliable foundation than static alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。