arXiv:2509.04512cs.CLcs.LG2025-09被引 1

小模型经微调可媲美大模型,适合隐私敏感的暖心对话场景

Scaling behavior of large language models in emotional safety classification across sizes and tasks

  • 用大模型生成情感重释数据,构建超1.5万样本的新心理安全分类数据集
  • 70B大模型在零样本下表现最优,但1B小模型微调后性能接近大模型
  • 1B模型仅需2GB显存,适合本地部署,保障用户隐私

理解大语言模型如何处理情绪敏感内容,对构建心理健康领域的安全可靠系统至关重要。本文研究了大模型在两类关键任务上的缩放行为:三分类情感安全判定(安全/不安全/边缘)与六类风险标签的多标签分类。为此,我们融合多个真人撰写的心理健康数据集(>1.5万样本),并利用ChatGPT生成情感重释提示进行数据增强。评估了四个LLaMA模型(1B、3B、8B、70B)在零样本、少样本和微调设置下的表现。结果表明,更大模型整体性能更强,尤其在复杂多标签分类和零样本场景中。然而,轻量级微调使1B模型在多个高数据类别中达到与大模型及BERT相当的性能,且推理时所需显存低于2GB。这说明小型化、本地部署的模型可作为敏感应用的可行替代方案,兼具情绪理解能力与安全边界控制。本研究为治疗型大模型应用及安全关键系统的可扩展对齐提供了重要启示。

原文摘要 · Abstract (English)

Understanding how large language models (LLMs) process emotionally sensitive content is critical for building safe and reliable systems, particularly in mental health contexts. We investigate the scaling behavior of LLMs on two key tasks: trinary classification of emotional safety (safe vs. unsafe vs. borderline) and multi-label classification using a six-category safety risk taxonomy. To support this, we construct a novel dataset by merging several human-authored mental health datasets (> 15K samples) and augmenting them with emotion re-interpretation prompts generated via ChatGPT. We evaluate four LLaMA models (1B, 3B, 8B, 70B) across zero-shot, few-shot, and fine-tuning settings. Our results show that larger LLMs achieve stronger average performance, particularly in nuanced multi-label classification and in zero-shot settings. However, lightweight fine-tuning allowed the 1B model to achieve performance comparable to larger models and BERT in several high-data categories, while requiring <2GB VRAM at inference. These findings suggest that smaller, on-device models can serve as viable, privacy-preserving alternatives for sensitive applications, offering the ability to interpret emotional context and maintain safe conversational boundaries. This work highlights key implications for therapeutic LLM applications and the scalable alignment of safety-critical systems.

情感安全小模型隐私保护心理健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。