构建首个罗马尼亚语作者画像数据集,利用Reddit社区特征推断作者背景。
Reddit is all you need: Authorship profiling for Romanian
- 基于Reddit子版块结构,从100多个社区提取用户写作样本。
- 构建2.3万+条标注数据,涵盖年龄、职业、兴趣等作者特征。
- 首次公开可用资源,推动小语种作者画像研究发展。
作者画像旨在通过文本识别作者特征,这一古老问题因自然语言处理技术发展而愈发重要。本文首次构建了一个罗马尼亚语短文本语料库,包含作者特征标签,数据源自社交平台Reddit。利用其基于主题的社区结构(子版块),结合用户发帖内容等线索,推断作者的年龄类别、职业状态、兴趣爱好及社会倾向等信息。最终获取超过2.3万条样本,来自100多个罗马尼亚语子版块。我们对数据集进行分析,并微调与评估大语言模型(LLMs),验证了基线性能,表明该领域仍需深入研究。所有资源均已公开。
原文摘要 · Abstract (English)
Authorship profiling is the process of identifying an author's characteristics based on their writings. This centuries old problem has become more intriguing especially with recent developments in Natural Language Processing (NLP). In this paper, we introduce a corpus of short texts in the Romanian language, annotated with certain author characteristic keywords; to our knowledge, the first of its kind. In order to do this, we exploit a social media platform called Reddit. We leverage its thematic community-based structure (subreddits structure), which offers information about the author's background. We infer an user's demographic and some broad personal traits, such as age category, employment status, interests, and social orientation based on the subreddit and other cues. We thus obtain a 23k+ samples corpus, extracted from 100+ Romanian subreddits. We analyse our dataset, and finally, we fine-tune and evaluate Large Language Models (LLMs) to prove baselines capabilities for authorship profiling using the corpus, indicating the need for further research in the field. We publicly release all our resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。