arXiv:2508.17576cs.CLcs.LG2025-08

用可解释的神经网络分析文本情感中关键词的因果效应。

CausalSent: Interpretable Sentiment Classification with RieszNet

  • 基于双头RieszNet架构,提升文本特征因果效应估计精度。
  • 在半合成IMDB数据上,效应估计误差降低2-3倍。
  • 发现'love'一词使正面情感概率提升2.9%,适合做可解释性研究。

尽管现代自然语言处理模型性能强劲,其决策过程仍如黑箱。为缩小这一差距,因果NLP结合因果推断与现代NLP模型,揭示文本特征的因果效应。我们复现并扩展Bansal等人的工作,聚焦于模型可解释性,提出基于RieszNet的双头神经网络架构,实现更精准的处理效应估计。我们的框架CausalSent在半合成的IMDB电影评论中表现优异,相比Bansal等人在合成Civil Comments数据上的均方误差(MAE),将效应估计误差降低2-3倍。通过集成验证过的模型,我们在观察性案例研究中发现,'love'一词的存在会使正面情感概率提升2.9%。

原文摘要 · Abstract (English)

Despite the overwhelming performance improvements offered by recent natural language processing (NLP) models, the decisions made by these models are largely a black box. Towards closing this gap, the field of causal NLP combines causal inference literature with modern NLP models to elucidate causal effects of text features. We replicate and extend Bansal et al's work on regularizing text classifiers to adhere to estimated effects, focusing instead on model interpretability. Specifically, we focus on developing a two-headed RieszNet-based neural network architecture which achieves better treatment effect estimation accuracy. Our framework, CausalSent, accurately predicts treatment effects in semi-synthetic IMDB movie reviews, reducing MAE of effect estimates by 2-3x compared to Bansal et al's MAE on synthetic Civil Comments data. With an ensemble of validated models, we perform an observational case study on the causal effect of the word "love" in IMDB movie reviews, finding that the presence of the word "love" causes a +2.9% increase in the probability of a positive sentiment.

可解释性因果推断情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。