用上下文感知的合成数据缓解心理防御机制分类的数据稀缺问题
VISHC at PsyDefDetect: Mitigating Data Scarcity in Psychological Defense Classification with Context-Aware Synthetic Augmentation

- 设计上下文感知的合成数据生成框架,结合临床特征增强生成真实性
- 在低资源场景下实现58.26%准确率和24.62%宏F1,显著优于基线
- 适合心理计算、临床自然语言处理领域研究者参考
心理防御机制(PDMs)是调节个体应对情绪困扰的无意识认知过程。从文本中自动分类PDMs具有临床价值,但受限于数据稀缺与类别不平衡,仅靠生成式增广无法解决。本文针对PsyDefDetect共享任务(BioNLP@ACL 2026),提出结合上下文感知合成增广与混合分类模型的方法。该模型融合上下文语言表示与150个标注的临床特征项。实验表明,提示词定义质量直接影响生成保真度与下游性能。所提方法超越DMRS Co-Pilot,在准确率上达58.26%(+40.25%),宏F1达24.62%(+15.99%),为低资源环境下心理导向的防御机制分类建立强基线。源码已公开:https://github.com/htdgv/CASA-PDC。
原文摘要 · Abstract (English)
Psychological defense mechanisms (PDMs) are unconscious cognitive processes that modulate how individuals perceive and respond to emotional distress. Automatically classifying PDMs from text is clinically valuable but severely hindered by data scarcity and class imbalance, challenges which generative augmentation alone cannot resolve without psychological grounding. In this work, we address these challenges in the PsyDefDetect shared task (BioNLP@ACL 2026) by proposing a context-aware synthetic augmentation framework combined with a hybrid classification model. Our hybrid model integrates contextual language representations with basic clinical features, along with 150 annotated defense items. Experiments demonstrate that definition quality in prompting directly governs generation fidelity and downstream performance. Our method surpasses DMRS Co-Pilot, reaching an accuracy of 58.26% (+40.25%) and a macro-F1 of 24.62% (+15.99%), thereby establishing a strong baseline for psychologically grounded defense mechanism classification in low-resource settings. Source code is available at: https://github.com/htdgv/CASA-PDC.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。