arXiv:2507.14693cs.CLcs.AI2025-07

构建土耳其语自杀意念数据集并评估标注可靠性与模型跨语言性能。

Rethinking Suicidal Ideation Detection: A Trustworthy Annotation Framework and Cross-Lingual Model Evaluation

  • 采用三人标注+大模型辅助的高效标注框架,构建土耳其语自杀意念语料库。
  • 发现主流模型在零样本迁移下表现可疑,标注一致性影响模型评估结果。
  • 推动心理健康NLP领域透明化标注与跨语言评估,适合研究心理安全与AI伦理者。

自杀意念检测对实时自杀预防至关重要,但其进展面临两大未充分探索的挑战:语言覆盖有限和标注实践不可靠。现有数据集多为英文,且高质量人工标注数据仍稀缺,许多研究依赖预标注数据而未检验其标注过程或标签可靠性。其他语言的数据集缺乏进一步限制了人工智能在全局自杀预防中的应用。本研究通过从社交媒体帖子构建新型土耳其语自杀意念语料库,并引入资源高效的三名人工标注员加两大型语言模型(LLMs)的标注框架。随后,基于八种预训练情感与情绪分类器进行迁移学习,双向评估该语料库与三个主流英文自杀意念检测数据集的标签可靠性与模型一致性。结果表明,需采用更严谨、包容多语言的标注与评估方法;同时揭示了主流模型在零样本迁移下的表现存疑。研究倡导心理健康自然语言处理中数据与模型构建的透明性,优先保障数据与模型的可靠性。

原文摘要 · Abstract (English)

Suicidal ideation detection is critical for real-time suicide prevention, yet its progress faces two under-explored challenges: limited language coverage and unreliable annotation practices. Most available datasets are in English, but even among these, high-quality, human-annotated data remains scarce. As a result, many studies rely on available pre-labeled datasets without examining their annotation process or label reliability. The lack of datasets in other languages further limits the global realization of suicide prevention via artificial intelligence (AI). In this study, we address one of these gaps by constructing a novel Turkish suicidal ideation corpus derived from social media posts and introducing a resource-efficient annotation framework involving three human annotators and two large language models (LLMs). We then address the remaining gaps by performing a bidirectional evaluation of label reliability and model consistency across this dataset and three popular English suicidal ideation detection datasets, using transfer learning through eight pre-trained sentiment and emotion classifiers. These transformers help assess annotation consistency and benchmark model performance against manually labeled data. Our findings underscore the need for more rigorous, language-inclusive approaches to annotation and evaluation in mental health natural language processing (NLP) while demonstrating the questionable performance of popular models with zero-shot transfer learning. We advocate for transparency in model training and dataset construction in mental health NLP, prioritizing data and model reliability.

心理健康跨语言标注可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。