arXiv:2505.12727cs.CLcs.CY2025-05ACL被引 9

构建心理疾病污名语料库,助力算法识别与消除偏见

What is Stigma Attributed to? A Theory-Grounded, Expert-Annotated Interview Corpus for Demystifying Mental-Health Stigma

  • 基于理论框架的人机对话访谈数据集,含4141个片段
  • 684名参与者具社会文化背景记录,提升标注可靠性
  • 公开数据集适合研究污名检测与干预策略

心理疾病污名仍是阻碍治疗与康复的普遍社会问题。现有用于训练神经模型细粒度分类污名的资源有限,主要依赖社交媒体或合成数据,缺乏理论基础。为弥补这一缺口,我们构建了一个由专家标注、理论指导的人机对话访谈语料库,包含来自684名参与者的4141个片段,且每位参与者均有明确的社会文化背景信息。实验对当前最先进的神经模型进行了基准测试,并实证揭示了污名检测面临的挑战。该数据集可推动计算方法在心理疾病污名检测、中和与对抗方面的研究。语料库已开源,地址为 https://github.com/HanMeng2004/Mental-Health-Stigma-Interview-Corpus。

原文摘要 · Abstract (English)

Mental-health stigma remains a pervasive social problem that hampers treatment-seeking and recovery. Existing resources for training neural models to finely classify such stigma are limited, relying primarily on social-media or synthetic data without theoretical underpinnings. To remedy this gap, we present an expert-annotated, theory-informed corpus of human-chatbot interviews, comprising 4,141 snippets from 684 participants with documented socio-cultural backgrounds. Our experiments benchmark state-of-the-art neural models and empirically unpack the challenges of stigma detection. This dataset can facilitate research on computationally detecting, neutralizing, and counteracting mental-health stigma. Our corpus is openly available at https://github.com/HanMeng2004/Mental-Health-Stigma-Interview-Corpus.

心理污名数据集自然语言处理社会影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。