arXiv:2601.00004cs.AIcs.CL2026-01被引 1

用大模型分析尼日利亚皮钦语语音,实现自动抑郁筛查。

Finetuning Large Language Models for Automated Depression Screening in Nigerian Pidgin English: GENSCORE Pilot Study

  • 用皮钦语对话数据微调大模型,适配本地语言和文化。
  • GPT-4.1在抑郁程度预测上准确率达94.5%。
  • 适合资源匮乏、多语言地区的心理健康筛查应用。

抑郁症是尼日利亚精神健康负担的主要因素,但因临床资源不足、污名化及语言障碍,筛查覆盖率低。传统量表如PHQ-9在高收入国家验证,对尼日利亚使用皮钦语的群体存在语言与文化不适应问题。本研究提出一种基于微调大语言模型(LLMs)的自动化抑郁筛查方法,针对尼日利亚青年(18-40岁)收集432条皮钦语语音回复,涵盖与PHQ-9项目一致的心理体验提问。经转录、预处理、语义标注、俚语解释与PHQ-9评分后,使用Phi-3-mini-4k-instruct、Gemma-3-4B-it和GPT-4.1三款模型进行微调。定量评估显示,GPT-4.1在PHQ-9严重程度预测中达到94.5%准确率,优于其他模型;定性评估中其回应最清晰、相关且文化贴合。该工作为在语言多样、资源有限环境中部署对话式心理健康工具提供了基础。

原文摘要 · Abstract (English)

Depression is a major contributor to the mental-health burden in Nigeria, yet screening coverage remains limited due to low access to clinicians, stigma, and language barriers. Traditional tools like the Patient Health Questionnaire-9 (PHQ-9) were validated in high-income countries but may be linguistically or culturally inaccessible for low- and middle-income countries and communities such as Nigeria where people communicate in Nigerian Pidgin and more than 520 local languages. This study presents a novel approach to automated depression screening using fine-tuned large language models (LLMs) adapted for conversational Nigerian Pidgin. We collected a dataset of 432 Pidgin-language audio responses from Nigerian young adults aged 18-40 to prompts assessing psychological experiences aligned with PHQ-9 items, performed transcription, rigorous preprocessing and annotation, including semantic labeling, slang and idiom interpretation, and PHQ-9 severity scoring. Three LLMs - Phi-3-mini-4k-instruct, Gemma-3-4B-it, and GPT-4.1 - were fine-tuned on this annotated dataset, and their performance was evaluated quantitatively (accuracy, precision and semantic alignment) and qualitatively (clarity, relevance, and cultural appropriateness). GPT-4.1 achieved the highest quantitative performance, with 94.5% accuracy in PHQ-9 severity scoring prediction, outperforming Gemma-3-4B-it and Phi-3-mini-4k-instruct. Qualitatively, GPT-4.1 also produced the most culturally appropriate, clear, and contextually relevant responses. AI-mediated depression screening for underserved Nigerian communities. This work provides a foundation for deploying conversational mental-health tools in linguistically diverse, resource-constrained environments.

抑郁筛查大模型皮钦语本土化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。