首个孟加拉语作者画像数据集,用机器学习分析社交文本识别人群属性。
BN-AuthProf: Benchmarking Machine Learning for Bangla Author Profiling on Social Media Texts
- 构建300人、3万条孟加拉语社交文本数据集,标注年龄与性别。
- 深度学习模型在年龄分类上准确率达91%,性别分类最高80%。
- 适用于隐私保护下的营销、司法与语言学研究,兼顾偏见控制。
作者画像旨在通过文本分析推断作者的性别、年龄等属性,随着社交媒体普及变得愈发重要。本文聚焦孟加拉语场景,提出并基准测试了机器学习方法在新构建的孟加拉语作者画像数据集 BN-AuthProf 上的表现。该数据集包含300名作者的30,131条社交文本,已匿名化处理以保障隐私。采用多种经典与深度学习模型进行评估:性别分类最佳准确率为80%(支持向量机,SVM),F1得分为0.756(多项式朴素贝叶斯,MNB);年龄分类中MNB取得最高准确率91%和F1分数0.905。研究表明,机器学习在孟加拉语作者画像任务中具有显著有效性,对市场营销、网络安全、法医语言学、教育及刑事侦查等领域具有实际应用价值,同时关注了隐私保护与模型偏差问题。
原文摘要 · Abstract (English)
Author profiling, the analysis of texts to uncover attributes such as gender and age of the author, has become essential with the widespread use of social media platforms. This paper focuses on author profiling in the Bangla language, aiming to extract valuable insights about anonymous authors based on their writing style on social media. The primary objective is to introduce and benchmark the performance of machine learning approaches on a newly created Bangla Author Profiling dataset, BN-AuthProf. The dataset comprises 30,131 social media posts from 300 authors, labeled by their age and gender. Authors' identities and sensitive information were anonymized to ensure privacy. Various classical machine learning and deep learning techniques were employed to evaluate the dataset. For gender classification, the best accuracy achieved was 80% using Support Vector Machine (SVM), while a Multinomial Naive Bayes (MNB) classifier achieved the best F1 score of 0.756. For age classification, MNB attained a maximum accuracy score of 91% with an F1 score of 0.905. This research highlights the effectiveness of machine learning in gender and age classification for Bangla author profiling, with practical implications spanning marketing, security, forensic linguistics, education, and criminal investigations, considering privacy and biases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。