小模型在心理健康理解上接近GPT-4,适合隐私敏感场景
Beyond Scale: Small Language Models are Comparable to GPT-4 in Mental Health Understanding
- 用零样本和少样本学习评估五款小模型与三款大模型表现
- 小模型在二分类任务上平均F1达0.64,仅比大模型低2%
- 少样本提示让小模型提升14.6%,适合快速定制化筛查工具
小语言模型(SLMs)作为敏感应用中的隐私保护替代方案,其内在理解能力是否可媲美大语言模型(LLMs)尚不明确。本文通过系统性评估六项心理健康理解任务,比较五款先进SLMs(Phi-3, Phi-3.5, Qwen2.5, Llama-3.2, Gemma2)与三款LLMs(GPT-4, FLAN-T5-XXL, Alpaca-7B)的性能。在零样本设置下,SLMs在二分类任务上的平均F1得分为0.64,仅比LLMs低2%;在多分类严重程度任务中,两类模型均出现超过30%的性能下降,表明临床理解挑战不依赖模型规模。少样本提示使SLMs性能最高提升14.6%,而LLMs增益更不稳定。结果表明,小模型具备有效处理敏感在线文本的能力,尤其在少量数据下快速适配,是可扩展的心理健康筛查工具的有力候选。
原文摘要 · Abstract (English)
The emergence of Small Language Models (SLMs) as privacy-preserving alternatives for sensitive applications raises a fundamental question about their inherent understanding capabilities compared to Large Language Models (LLMs). This paper investigates the mental health understanding capabilities of current SLMs through systematic evaluation across diverse classification tasks. Employing zero-shot and few-shot learning paradigms, we benchmark their performance against established LLM baselines to elucidate their relative strengths and limitations in this critical domain. We assess five state-of-the-art SLMs (Phi-3, Phi-3.5, Qwen2.5, Llama-3.2, Gemma2) against three LLMs (GPT-4, FLAN-T5-XXL, Alpaca-7B) on six mental health understanding tasks. Our findings reveal that SLMs achieve mean performance within 2\% of LLMs on binary classification tasks (F1 scores of 0.64 vs 0.66 in zero-shot settings), demonstrating notable competence despite orders of magnitude fewer parameters. Both model categories experience similar degradation on multi-class severity tasks (a drop of over 30\%), suggesting that nuanced clinical understanding challenges transcend model scale. Few-shot prompting provides substantial improvements for SLMs (up to 14.6\%), while LLM gains are more variable. Our work highlights the potential of SLMs in mental health understanding, showing they can be effective privacy-preserving tools for analyzing sensitive online text data. In particular, their ability to quickly adapt and specialize with minimal data through few-shot learning positions them as promising candidates for scalable mental health screening tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。