arXiv:2504.00265cs.CL2025-04被引 1

研究不同语言摘要对情感分析的影响,发现提取式更保情感

Multilingual Sentiment Analysis of Summarized Texts: A Cross-Language Study of Text Shortening Effects

  • 对比提取与抽象摘要在8种语言中的情感保留效果
  • 复杂形态语言(如芬兰语、匈牙利语)准确率下降超30%
  • 适合多语言舆情监控和跨语言分析场景

摘要显著影响具有不同形态结构的语言的情感分析。本研究考察了英语、德语、法语、西班牙语、意大利语、芬兰语、匈牙利语和阿拉伯语中提取式与抽象式摘要对情感分类的影响。我们使用多语言Transformer模型(mBERT、XLM-RoBERTa、T5、BART)及语言特定模型(FinBERT、AraBERT)评估摘要后的情感变化。结果表明,提取式摘要更能保持情感倾向,尤其在形态复杂的语言中;而抽象式摘要虽提升可读性,但引入情感偏差,降低分类准确率。具有丰富屈折形态的语言(如芬兰语、匈牙利语、阿拉伯语)准确率下降幅度显著高于英语或德语。研究强调情感分析需考虑语言特性,并提出兼顾可读性与情感保留的混合摘要方法。研究成果对社交媒体监测、市场分析及跨语言意见挖掘等应用具有重要价值。

原文摘要 · Abstract (English)

Summarization significantly impacts sentiment analysis across languages with diverse morphologies. This study examines extractive and abstractive summarization effects on sentiment classification in English, German, French, Spanish, Italian, Finnish, Hungarian, and Arabic. We assess sentiment shifts post-summarization using multilingual transformers (mBERT, XLM-RoBERTa, T5, and BART) and language-specific models (FinBERT, AraBERT). Results show extractive summarization better preserves sentiment, especially in morphologically complex languages, while abstractive summarization improves readability but introduces sentiment distortion, affecting sentiment accuracy. Languages with rich inflectional morphology, such as Finnish, Hungarian, and Arabic, experience greater accuracy drops than English or German. Findings emphasize the need for language-specific adaptations in sentiment analysis and propose a hybrid summarization approach balancing readability and sentiment preservation. These insights benefit multilingual sentiment applications, including social media monitoring, market analysis, and cross-lingual opinion mining.

情感分析多语言摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。