让大模型更共情:用人类协作提升低资源语言心理辅导质量
Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

- 设计角色扮演反思链提示框架,结合专家案例与自我反思生成共情回复
- 在625个真实案例上测试,新方法使响应更贴近专业心理咨询水平
- 适合研究跨文化心理AI、人机协同咨询或低资源语言NLP的开发者
尽管大语言模型取得进展,其在低资源语言中生成共情心理辅导回应的能力仍待探索。为此,我们从三个来源收集了625个真实心理求助案例:(1)公开的Facebook帖子,(2)孟加拉国电视节目《Ami Akhon Ki Korbo》的访谈记录,(3)匿名学生问卷中反映的多样化情绪与心理挑战。基于这些案例,我们构建了一个评估语料库,包含持证临床心理学家撰写的建议以及GPT-4o Mini、Claude 4.5 Haiku和Gemini 2.5 Pro三款主流商用LLM生成的回复。我们提出针对特定任务的“角色扮演反思链式咨询框架”(RP-RCAF),通过专家少样本示例与结构化自我反思,以共情顾问角色生成支持性、文化敏感且符合伦理的建议。同时引入基于Grok 4的评估与评分框架(G-REFS),融合自动化评估与心理学专家验证,涵盖情感敏感度、文化适宜性、语言清晰度和伦理合理性。实验表明,RP-RCAF在所有测试模型上均优于传统提示策略,生成结果更接近专业心理辅导标准。
原文摘要 · Abstract (English)
Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program "Ami Akhon Ki Korbo", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。