用多维极性场分析话语转变,可解释语义迁移方向与程度。
TOPol: Capturing and Explaining Multidimensional Semantic Polarity Fields and Vectors
- 基于大模型嵌入+聚类分割,构建可解释的多维极性向量场。
- 在美联储讲话和亚马逊评论中均捕捉到情感与非情感极性变化。
- 适合需要理解话语转变机制的研究者或政策分析人员。
传统计算语言学将语义极性视为单维尺度,忽略了语言的多维结构。本文提出TOPol(主题-方向极性)框架,一种半监督方法,在人类在环定义的上下文边界(CBs)下重建并解析多维叙事极性场。该框架使用基于Transformer的大语言模型(tLLM)嵌入文档,通过邻域调优的UMAP投影,并采用Leiden算法进行主题分割。给定两个论述范式A与B之间的CB,TOPol计算对应主题边界中心间的有向向量,生成量化范式转换期间细粒度语义位移的极性场。该向量表示可用于评估CB质量并检测极性变化,指导人类在环的CB优化。为解释极性向量,tLLM对比其端点并生成具有估计覆盖范围的对比标签。稳健性分析显示,仅CB定义(人类在环可调节的主要参数)显著影响结果,证实方法稳定性。我们在两个语料库上评估TOPol:(i) 美联储在宏观经济转折点附近的演讲,捕捉非情感语义转移;(ii) 亚马逊产品评论按评分分层,其中情感极性与NRC情感值一致。结果表明,TOPol持续捕捉情感与非情感极性转换,提供一种可扩展、泛化性强且可解释的上下文敏感多维话语分析框架。
原文摘要 · Abstract (English)
Traditional approaches to semantic polarity in computational linguistics treat sentiment as a unidimensional scale, overlooking the multidimensional structure of language. This work introduces TOPol (Topic-Orientation POLarity), a semi-unsupervised framework for reconstructing and interpreting multidimensional narrative polarity fields under human-on-the-loop (HoTL) defined contextual boundaries (CBs). The framework embeds documents using a transformer-based large language model (tLLM), applies neighbor-tuned UMAP projection, and segments topics via Leiden partitioning. Given a CB between discourse regimes A and B, TOPol computes directional vectors between corresponding topic-boundary centroids, yielding a polarity field that quantifies fine-grained semantic displacement during regime shifts. This vectorial representation enables assessing CB quality and detecting polarity changes, guiding HoTL CB refinement. To interpret identified polarity vectors, the tLLM compares their extreme points and produces contrastive labels with estimated coverage. Robustness analyses show that only CB definitions (the main HoTL-tunable parameter) significantly affect results, confirming methodological stability. We evaluate TOPol on two corpora: (i) U.S. Central Bank speeches around a macroeconomic breakpoint, capturing non-affective semantic shifts, and (ii) Amazon product reviews across rating strata, where affective polarity aligns with NRC valence. Results demonstrate that TOPol consistently captures both affective and non-affective polarity transitions, providing a scalable, generalizable, and interpretable framework for context-sensitive multidimensional discourse analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。