首个跨语言自杀倾向文本检测模型,支持六种语言实时识别风险内容。
The First Multilingual Model For The Detection of Suicide Texts
- 基于mT5等Transformer架构构建多语言检测模型
- mT5在六语言上实现超85%的F1分数,表现最优
- 适用于心理健康监测与跨文化干预,适合安全与伦理研究者
自杀意念是全球影响数百万人的重大健康问题。社交媒体通过用户情绪表达提供了相关线索。本文提出首个多语言模型,利用mBERT、XML-R和mT5等Transformer架构,在西班牙语、英语、德语、加泰罗尼亚语、葡萄牙语和意大利语共六种语言中检测自杀文本。使用SeamlessM4T将西班牙语自杀意念推文数据集翻译至其他五种语言,每种模型均在此多语言数据上微调并评估分类性能。结果显示,mT5整体表现最佳,各语言F1分数均超过85%,凸显其跨语言迁移学习能力。英文与西班牙语翻译在困惑度上表现优异,证明质量较高。本研究强调在开发自动化多语言工具时需重视语言多样性。局限性包括翻译中的语义保真度问题及伦理考量,提示未来需开展人机协同评估。
原文摘要 · Abstract (English)
Suicidal ideation is a serious health problem affecting millions of people worldwide. Social networks provide information about these mental health problems through users' emotional expressions. We propose a multilingual model leveraging transformer architectures like mBERT, XML-R, and mT5 to detect suicidal text across posts in six languages - Spanish, English, German, Catalan, Portuguese and Italian. A Spanish suicide ideation tweet dataset was translated into five other languages using SeamlessM4T. Each model was fine-tuned on this multilingual data and evaluated across classification metrics. Results showed mT5 achieving the best performance overall with F1 scores above 85%, highlighting capabilities for cross-lingual transfer learning. The English and Spanish translations also displayed high quality based on perplexity. Our exploration underscores the importance of considering linguistic diversity in developing automated multilingual tools to identify suicidal risk. Limitations exist around semantic fidelity in translations and ethical implications which provide guidance for future human-in-the-loop evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。