从混乱社交媒体文本中恢复丢失的emoji,提升数据可用性。
Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and A Case Study
- 提出三步逆向工程法,从乱码文本中提取emoji
- 在50万条新冠相关推文中找回15万余个emoji
- 适合研究社交情绪、数据清洗与文本可读性分析者
表情符号广泛用于社交媒体,但在噪声或乱码文本中常被丢失,影响数据分析与机器学习。传统预处理方法建议删除此类文本,却可能丢弃表情符号及其语境意义。本文提出一种三步逆向工程方法,用于从社交媒体帖子的乱码文本中恢复表情符号,并识别其生成原因。为评估效果,该方法应用于包含509,248条关于猴痘疫情的推文数据集,该数据集曾被约30篇先期研究引用但未能提取表情符号。本方法成功从76,914条推文中恢复157,748个表情符号。通过Flesch Reading Ease、Flesch-Kincaid Grade Level、Coleman-Liau Index、Automated Readability Index、Dale-Chall Readability Score、Text Standard及Reading Time等指标,验证了文本可读性与连贯性的提升。同时分析了各表情符号使用频率及模式,结果已呈现。
原文摘要 · Abstract (English)
Emojis are widely used across social media platforms but are often lost in noisy or garbled text, posing challenges for data analysis and machine learning. Conventional preprocessing approaches recommend removing such text, risking the loss of emojis and their contextual meaning. This paper proposes a three-step reverse-engineering methodology to retrieve emojis from garbled text in social media posts. The methodology also identifies reasons for the generation of such text during social media data mining. To evaluate its effectiveness, the approach was applied to 509,248 Tweets about the Mpox outbreak, a dataset referenced in about 30 prior works that failed to retrieve emojis from garbled text. Our method retrieved 157,748 emojis from 76,914 Tweets. Improvements in text readability and coherence were demonstrated through metrics such as Flesch Reading Ease, Flesch-Kincaid Grade Level, Coleman-Liau Index, Automated Readability Index, Dale-Chall Readability Score, Text Standard, and Reading Time. Additionally, the frequency of individual emojis and their patterns of usage in these Tweets were analyzed, and the results are presented.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。