小模型比大模型更擅长识别精神分裂症的思维混乱症状。
Bigger But Not Better: Small Neural Language Models Outperform Large Language Models in Detection of Thought Disorder
- 用滑动窗口困惑度检测思维混乱,小模型表现优于大模型。
- 模型过大反而降低检测敏感性,存在最佳尺寸阈值。
- 适合临床筛查,兼顾隐私、成本与部署便利性。
思维紊乱是精神分裂症谱系障碍的关键诊断指标。近期研究发现,临床对思维紊乱严重程度的评估与大型语言模型(LLMs)预测话语难度的相关性显著。然而,LLMs存在隐私风险、计算与资金成本高、训练数据不透明等问题,限制其临床应用。本文探究了小型神经语言模型是否可作为有效替代方案,采用与大模型相同的滑动窗口困惑度测量方法检测阳性形式思维紊乱。结果令人意外:小模型对与思维紊乱相关的语言差异更敏感,模型规模和上下文长度超过一定阈值后,检测能力反而下降,挑战了“越大越好”的普遍认知。该结论在精神症状患者的声音日记和临床访谈语料中均得到验证,表明此类方法为开发高效、低成本、隐私友好的筛查工具提供了新方向,适用于临床及自然场景部署。
原文摘要 · Abstract (English)
Disorganized thinking is a key diagnostic indicator of schizophrenia-spectrum disorders. Recently, clinical estimates of the severity of disorganized thinking have been shown to correlate with measures of how difficult speech transcripts would be for large language models (LLMs) to predict. However, LLMs' deployment challenges -- including privacy concerns, computational and financial costs, and lack of transparency of training data -- limit their clinical utility. We investigate whether smaller neural language models can serve as effective alternatives for detecting positive formal thought disorder, using the same sliding window based perplexity measurements that proved effective with larger models. Surprisingly, our results show that smaller models are more sensitive to linguistic differences associated with formal thought disorder than their larger counterparts. Detection capability declines beyond a certain model size and context length, challenging the common assumption of ``bigger is better'' for LLM-based applications. Our findings generalize across audio diaries and clinical interview speech samples from individuals with psychotic symptoms, suggesting a promising direction for developing efficient, cost-effective, and privacy-preserving screening tools that can be deployed in both clinical and naturalistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。