arXiv:2606.00647cs.CLcs.AI2026-06中稿 · PsyDefDetect, a sh…被引 1

针对心理防御机制分类中的类别不平衡问题,改进Qwen3-8B模型表现。

LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification

论文配图:LinguIUTics at PsyDefDetect: Iterative Imbalance-Aware Fine-tuning of Qwen3-8B for Psychological Defense Mechanism Classification
图 1 · 摘自论文原文
  • 采用分组分层交叉验证与轮转词汇增强,缓解少数类数据不足
  • 通过后处理调整逻辑偏置并集成多模型,使关键类别F1达0.797
  • 适合临床对话分析与罕见类识别任务的研究者参考

在对话文本中检测心理防御机制仍是临床NLP中的难题。针对PsyDefDetect 2026共享任务(九类话语分类,以宏平均F1为评估指标),我们团队LinguIUTics在官方正类排行榜上取得0.3917的宏平均F1,位列21支参赛队伍中的第4名,相比Ministral-8B基线(31.48宏平均F1)提升7.7个百分点(相对提升24.4%)。基于BERT的编码器和零样本大模型在稀有类别上表现不佳,主要因严重类别不平衡。因此我们采用QLoRA对Qwen3-8B进行迭代式不平衡感知微调。核心策略包括:分组分层交叉验证(防止信息泄露)、少数类轮转词汇增强、以及结合逻辑偏置调节与集成融合的后处理流程。三者协同显著缩小了验证集与排行榜差距,大幅提升少数类召回率,使关键类别“不确定”(第8级)的F1从接近零跃升至0.797。

原文摘要 · Abstract (English)

Detecting psychological defense mechanisms in conversational text remains a challenging clinical NLP problem. For the PsyDefDetect 2026 shared task (nine-class utterance classification evaluated via macro F1), our team LinguIUTics achieves a macro F1-score of 0.3917 on the official positive-class leaderboard, ranking 4th out of 21 registered teams and improving over the Ministral-8B task baseline (31.48 macro F1) by 7.7 absolute points (24.4 percent relative). BERT-family encoders and zero-shot LLMs proved ineffective on rare classes due to severe class imbalance, leading us to QLoRA fine-tuning of Qwen3-8B. We leverage three key strategies: grouped stratified cross-validation (preventing leakage), minority-class round-robin lexical augmentation, and a post-processing pipeline with logit bias tuning and ensemble blending. Together, these components close much of the validation-to-leaderboard gap and substantially improve minority-class recall, driving the critical "Unclear" class (Level 8) from near-zero performance to an F1 score of 0.797.

心理分析大模型微调类别不平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。