arXiv:2602.18443cs.HCcs.AI2026-02中稿 · the Springer CCIS …被引 1

评估大模型生成心理辅导邮件主题行的实用性与伦理风险。

From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications

  • 用分级评估法先分类再排序,提升评价可操作性。
  • 德语微调后模型表现显著优于通用模型,隐私开源模型性能接近商用。
  • 揭示心理AI部署中的隐私、偏见与责任问题,适合关注AI伦理的研究者。

心理社交在线咨询常因主题行过于泛化而难以高效分诊。本研究通过层级评估方法,对十一款大语言模型生成的六字德语咨询邮件主题行进行评价:先分类,再同类别内排序,实现可管理的评估。九名评估者(心理咨询师与AI系统)使用克里彭多夫α系数、斯皮尔曼等级相关系数、皮尔逊相关系数及肯德尔等级相关系数进行分析。结果表明,专有服务与注重隐私的开源模型间存在性能权衡,德语微调显著提升效果。研究还探讨了心理健康AI应用中的隐私、偏见与问责等关键伦理议题。

原文摘要 · Abstract (English)

Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates eleven large language models generating six-word subject lines for German counselling emails through hierarchical assessment - first categorising outputs, then ranking within categories to enable manageable evaluation. Nine assessors (counselling professionals and AI systems) enable analysis via Krippendorff's $α$, Spearman's $ρ$, Pearson's $r$ and Kendall's $τ$. Results reveal performance trade-offs between proprietary services and privacy-preserving open-source alternatives, with German fine-tuning consistently improving performance. The study addresses critical ethical considerations for mental health AI deployment including privacy, bias and accountability.

心理健康大模型评估伦理风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。