arXiv:2603.20222cs.CL2026-03被引 1

通过提取语言特征提升情绪识别准确率,效果稳定且可解释。

Linguistic Signatures for Enhanced Emotion Detection

  • 从13个英文数据集提取情绪特异性语言信号
  • 融合高阶语言特征后,宏F1提升最高达2.4点
  • 适合关注模型可解释性与鲁棒性的研究者

情绪检测是自然语言处理的核心问题,近年来得益于基于Transformer的模型和成熟数据集的进步。然而,关于情绪在不同语料和标签中如何表达的规律仍不清晰。本研究探讨语言特征是否能作为情绪识别的可靠可解释信号。我们从13个英文数据集提取情绪特异性语言签名,并评估将这些特征融入Transformer模型的效果。基于RoBERTa的模型在加入高层次语言特征后,在GoEmotions基准上宏F1最高提升2.4点,表明显式词汇线索可补充神经表征,提升情绪分类预测的鲁棒性。

原文摘要 · Abstract (English)

Emotion detection is a central problem in NLP, with recent progress driven by transformer-based models trained on established datasets. However, little is known about the linguistic regularities that characterize how emotions are expressed across different corpora and labels. This study examines whether linguistic features can serve as reliable interpretable signals for emotion recognition in text. We extract emotion-specific linguistic signatures from 13 English datasets and evaluate how incorporating these features into transformer models impacts performance. Our RoBERTa-based models enriched with high level linguistic features achieve consistent performance gains of up to +2.4 macro F1 on the GoEmotions benchmark, showing that explicit lexical cues can complement neural representations and improve robustness in predicting emotion categories.

情绪识别可解释性语言特征Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。