arXiv:2508.03112cs.CL2025-08

跨语言新闻对比中挖掘情感差异,发现同源媒体更一致。

Cross-lingual Opinions and Emotions Mining in Comparable Documents

  • 用双语词典标注跨语言新闻的情感与情绪,无需机器翻译。
  • 同源媒体(如BBC、阿联酋)文章情感一致性高,异源则差异明显。
  • 方法可推广至其他语言对,适合多语言舆情分析研究者。

可比文本是主题一致但非直译的多语言文档,有助于理解话题在不同语言中的表达差异。本研究聚焦英语-阿拉伯语可比文档中的情感与情绪差异。首先对文本进行情感与情绪标签标注,采用跨语言方法识别观点类别(主观/客观),避免依赖机器翻译。为标注情绪(愤怒、厌恶、恐惧、喜悦、悲伤、惊讶),将英文WordNet-Affect(WNA)词典手动翻译为阿拉伯语,构建双语情绪词典,用于标注可比语料库。随后使用统计方法评估每对源-目标文档间的情感与情绪一致性。该比较在文档来源不同时尤为相关,此前文献未充分探讨此问题。研究涵盖来自Euronews、BBC和Al-Jazeera(JSC)的英阿文档对。结果显示,同源媒体的文章情感与情绪标注高度一致,而异源媒体则显著分歧。所提方法具有语言无关性,可推广至其他语言对。

原文摘要 · Abstract (English)

Comparable texts are topic-aligned documents in multiple languages that are not direct translations. They are valuable for understanding how a topic is discussed across languages. This research studies differences in sentiments and emotions across English-Arabic comparable documents. First, texts are annotated with sentiment and emotion labels. We apply a cross-lingual method to label documents with opinion classes (subjective/objective), avoiding reliance on machine translation. To annotate with emotions (anger, disgust, fear, joy, sadness, surprise), we manually translate the English WordNet-Affect (WNA) lexicon into Arabic, creating bilingual emotion lexicons used to label the comparable corpora. We then apply a statistical measure to assess the agreement of sentiments and emotions in each source-target document pair. This comparison is especially relevant when the documents originate from different sources. To our knowledge, this aspect has not been explored in prior literature. Our study includes English-Arabic document pairs from Euronews, BBC, and Al-Jazeera (JSC). Results show that sentiment and emotion annotations align when articles come from the same news agency and diverge when they come from different ones. The proposed method is language-independent and generalizable to other language pairs.

跨语言情感分析新闻挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。