arXiv:2510.11739cs.SIcs.AI2025-10

用粉丝推文分析乌尔都语名人,预测性别、年龄等特征。

Celebrity Profiling on Short Urdu Text using Twitter Followers' Feed

  • 基于粉丝推文的文本特征,用多种机器学习模型预测名人属性。
  • 性别预测准确率达65%,在低资源语言中表现良好。
  • 为小语种社交媒体分析提供新思路,适合做社会计算研究者参考。

社交媒体已成为数字时代的重要组成部分,是交流、互动和信息共享的平台。名人是其中最活跃的用户之一,常通过在线帖子展现个人与职业生活。推特等平台为分析语言与行为模式提供了机会,以理解人口与社会规律。由于粉丝常与关注的名人具有相似的语言特质与兴趣,其文本数据可用于预测名人属性。然而,现有研究多集中于英语及其他高资源语言,乌尔都语研究仍属空白。本研究应用现代机器学习与深度学习技术解决乌尔都语名人的画像问题。收集并预处理了来自南亚名人粉丝的短篇乌尔都语推文数据集,训练并比较了逻辑回归、支持向量机、随机森林、卷积神经网络与长短期记忆网络等多种算法。采用准确率、精确率、召回率、F1分数与累积排名(cRank)进行评估。最佳结果为性别预测,cRank达0.65,准确率亦为0.65;年龄、职业与声望预测取得中等效果。结果表明,利用粉丝语言特征,结合机器学习与神经网络方法,可在乌尔都语这一低资源语言中有效实现名人属性预测。

原文摘要 · Abstract (English)

Social media has become an essential part of the digital age, serving as a platform for communication, interaction, and information sharing. Celebrities are among the most active users and often reveal aspects of their personal and professional lives through online posts. Platforms such as Twitter provide an opportunity to analyze language and behavior for understanding demographic and social patterns. Since followers frequently share linguistic traits and interests with the celebrities they follow, textual data from followers can be used to predict celebrity demographics. However, most existing research in this field has focused on English and other high-resource languages, leaving Urdu largely unexplored. This study applies modern machine learning and deep learning techniques to the problem of celebrity profiling in Urdu. A dataset of short Urdu tweets from followers of subcontinent celebrities was collected and preprocessed. Multiple algorithms were trained and compared, including Logistic Regression, Support Vector Machines, Random Forests, Convolutional Neural Networks, and Long Short-Term Memory networks. The models were evaluated using accuracy, precision, recall, F1-score, and cumulative rank (cRank). The best performance was achieved for gender prediction with a cRank of 0.65 and an accuracy of 0.65, followed by moderate results for age, profession, and fame prediction. These results demonstrate that follower-based linguistic features can be effectively leveraged using machine learning and neural approaches for demographic prediction in Urdu, a low-resource language.

名人画像乌尔都语社交媒体分析低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。