首个德语年龄标注毒评数据集,揭示不同年龄段网络言辞差异。
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
- 结合人工与大模型标注,构建3万+条德语社交评论数据
- 16.7%评论被标为问题内容,年轻人多用表达性语言
- 适合研究年龄差异、开发更公平的平台审核系统
现有毒评数据集缺乏人口统计背景,限制了对不同年龄群体网络交流方式的理解。本研究与德国公共广播机构funk合作,首次发布大规模德语毒评数据集,包含3,024条人工标注和30,024条大语言模型标注的匿名评论,来源为Instagram、TikTok和YouTube。通过预定义毒评关键词筛选,最终16.7%的评论被标记为问题内容。标注流程融合人工专家判断与先进语言模型,识别出辱骂、虚假信息、对广播费的批评等关键类别。数据显示,年轻用户更倾向使用表达性语言,而年长用户更多涉及虚假信息传播与贬低性言论。该资源为跨年龄语言差异研究提供新可能,支持开发更具公平性与年龄敏感性的内容审核系统。
原文摘要 · Abstract (English)
A lack of demographic context in existing toxic speech datasets limits our understanding of how different age groups communicate online. In collaboration with funk, a German public service content network, this research introduces the first large-scale German dataset annotated for toxicity and enriched with platform-provided age estimates. The dataset includes 3,024 human-annotated and 30,024 LLM-annotated anonymized comments from Instagram, TikTok, and YouTube. To ensure relevance, comments were consolidated using predefined toxic keywords, resulting in 16.7\% labeled as problematic. The annotation pipeline combined human expertise with state-of-the-art language models, identifying key categories such as insults, disinformation, and criticism of broadcasting fees. The dataset reveals age-based differences in toxic speech patterns, with younger users favoring expressive language and older users more often engaging in disinformation and devaluation. This resource provides new opportunities for studying linguistic variation across demographics and supports the development of more equitable and age-aware content moderation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。