arXiv:2603.13038cs.CL2026-03ACL被引 2

提出一种PCA降维选择新方法,让文本语义分析更稳定可解释。

Interpretable Semantic Gradients in SSD: A PCA Sweep Approach and a Case Study on AI Discourse

  • 通过系统扫描不同主成分数,自动选出最优降维维度
  • 发现自恋中的钦佩维度与对AI的乐观协作态度相关
  • 适合关注心理特质与语言意义关联的研究者使用

监督语义差异(SSD)是一种混合定量-解释性方法,通过在嵌入空间中估计语义梯度,并结合聚类与文本检索来解读其端点,分析文本意义如何随个体差异变量变化。现有方法在回归前使用PCA,但缺乏系统性的主成分数量选择标准,导致分析中存在可避免的研究者自由度。本文提出一种PCA扫视法,将降维维度选择视为表示能力、梯度可解释性及邻近K值间稳定性三者的联合准则。以普罗利菲克平台参与者撰写的关于人工智能的短帖为语料,同时采集其自恋中的钦佩与竞争量表得分,进行案例研究。结果显示,经扫视法确定的降维方案能生成稳定且可解释的钦佩相关语义梯度,对比出对AI持乐观合作态度与不信任讽刺言论;而竞争维度未出现显著对应关系。反事实分析表明,采用高维主成分的启发式方法会产生松散、结构弱的聚类,进一步验证扫视法的优势。该案例说明,该方法在限制研究者自由度的同时,保持了SSD的解释目标,支持透明且具有心理学意义的内涵意义分析。

原文摘要 · Abstract (English)

Supervised Semantic Differential (SSD) is a mixed quantitative-interpretive method that models how text meaning varies with continuous individual-difference variables by estimating a semantic gradient in an embedding space and interpreting its poles through clustering and text retrieval. SSD applies PCA before regression, but currently no systematic method exists for choosing the number of retained components, introducing avoidable researcher degrees of freedom in the analysis pipeline. We propose a PCA sweep procedure that treats dimensionality selection as a joint criterion over representation capacity, gradient interpretability, and stability across nearby values of K. We illustrate the method on a corpus of short posts about artificial intelligence written by Prolific participants who also completed Admiration and Rivalry narcissism scales. The sweep yields a stable, interpretable Admiration-related gradient contrasting optimistic, collaborative framings of AI with distrustful and derisive discourse, while no robust alignment emerges for Rivalry. We also show that a counterfactual using a high-PCA dimension solution heuristic produces diffuse, weakly structured clusters instead, reinforcing the value of the sweep-based choice of K. The case study shows how the PCA sweep constrains researcher degrees of freedom while preserving SSD's interpretive aims, supporting transparent and psychologically meaningful analyses of connotative meaning.

语义分析降维方法心理测量可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。