arXiv:2608.00497cs.CL2026-08被引 1

构建超5万条俄语社交文本数据集,助力识别自杀预警与反自杀信号。

The methodology of Constructing the Large-Scale Dataset for Detecting Presuicidal and Anti-Suicidal Signals in Social Media Texts in Russian

  • 设计全流程标注方法,包含指令制定、标注表创建与多轮校验
  • 建成包含5万余条俄语社交文本的大规模数据集,含预警与反向信号标注
  • 开源数据集与代码,适合心理健康与社会计算研究者使用

自杀是受心理状态误导的可怕行为,全球多国均面临此问题,俄罗斯亦然。幸运的是,部分有自杀倾向者会在社交媒体上表达内心挣扎,为及时干预提供可能。然而,这些关键文本常被大量无关内容淹没,严重延迟风险判断。为此,本文提出一套完整的数据集构建方法,涵盖标注指南与标签表设计、标注流程、验证及后期修正机制。基于该方法,我们收集并标注了超过5万条俄语社交媒体文本,形成大规模数据集,并提供统计分析与标注常见问题说明。此外,我们开展了基础分类模型实验,评估不同标注层级下的性能表现。所有数据集、代码与材料均已公开可用。

原文摘要 · Abstract (English)

The suicide is a terrifying act of a person who is misled by his own mental state. This problem arises across many countries. Sadly, Russia also has quite high number of persons who committed suicide. Luckily, a subset of these people writes their struggles in social media, allowing a way to find them and help. However, these valuable texts disappearing in many irrelevant texts which is considerably slowing down the decision process about person's suicidal risk. To tackle this problem, in this work we have presented a detailed methodology of building the dataset for detecting texts that describe presuicidal and anti-suicidal signals. This methodology describes the process of instruction and class table creation, the process of annotation, verification and post-annotation correction. Guiding by this methodology, we collect and annotate a large-scale Russian dataset with more than 50 thousand texts from social media. We provide a count statistic of the dataset as well as common problems in annotation. We also conduct basic experiments of building the classification models to show the on go performance on different levels of annotation. Furthermore, we make the dataset, code and all materials publicly available.

自杀检测俄语数据集社交媒体分析心理健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。