arXiv:2504.07245cs.AIcs.LG2025-04被引 2

用教师学生架构和新损失函数提升心理健康文本分类效果

A new training approach for text classification in Mental Health: LatentGLoss

  • 设计双模型结构,通过修改损失函数融合教师输出与隐层特征
  • 在心理健康数据集上,新方法显著优于传统模型与标准蒸馏
  • 适合关注医疗文本分类与模型压缩的科研人员使用

本研究提出一种多阶段心理健康分类方法,结合传统机器学习、深度神经网络与基于Transformer的模型。构建了一个新的数据集用于评估不同方法的性能,从基础分类器逐步推进至神经网络。评估了循环神经网络(如LSTM、GRU)对序列模式建模的有效性,并微调BERT等Transformer模型以检验上下文嵌入的影响。核心贡献在于一种新型训练策略:采用教师-学生双模型架构,不依赖软标签传递,而是通过修改损失函数,使教师模型的输出与隐层表示共同指导学生模型训练。实验结果表明,各阶段模型均有效,且所提损失函数与师生交互机制显著提升了模型在心理健康预测任务中的学习能力。

原文摘要 · Abstract (English)

This study presents a multi-stage approach to mental health classification by leveraging traditional machine learning algorithms, deep learning architectures, and transformer-based models. A novel data set was curated and utilized to evaluate the performance of various methods, starting with conventional classifiers and advancing through neural networks. To broaden the architectural scope, recurrent neural networks (RNNs) such as LSTM and GRU were also evaluated to explore their effectiveness in modeling sequential patterns in the data. Subsequently, transformer models such as BERT were fine-tuned to assess the impact of contextual embeddings in this domain. Beyond these baseline evaluations, the core contribution of this study lies in a novel training strategy involving a dual-model architecture composed of a teacher and a student network. Unlike standard distillation techniques, this method does not rely on soft label transfer; instead, it facilitates information flow through both the teacher model's output and its latent representations by modifying the loss function. The experimental results highlight the effectiveness of each modeling stage and demonstrate that the proposed loss function and teacher-student interaction significantly enhance the model's learning capacity in mental health prediction tasks.

心理健康文本分类知识蒸馏Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。