用多智能体辩论与缺失推理提升心理防御机制识别准确率
UTS at PsyDefDetect: Multi-Agent Councils and Absence-Based Reasoning for Defense Mechanism Classification
- 以情绪-认知融合谱为线索,通过提示规则捕捉防御机制的缺失特征
- 无微调下达F1 0.382(Top 5),结合纠错模型后提升至0.406
- 适合心理分析、人机对话评估等需细粒度情绪理解的场景
本文介绍在情感支持对话中分类心理防御机制的系统,基于防御机制评分量表(DMRS)参赛,获64支队伍中第二名(F1 0.406)。核心洞察是:防御机制由缺失构成——缺乏情感、阻断认知、否认现实。我们将其编码为提示层级的临床规则,沿情绪-认知整合谱推进,带来最大单点增益(+11.4pp F1)。系统采用多阶段协商式多智能体架构,由Gemini 2.5代理担任类属倡导者,评估证据强度而非投票,实现无微调下F1 0.382(独立排名前五)。但发现少数类预测易错:59%-80%稳定少数类预测错误,源于“L7吸引子”——情感内容默认归入多数类。通过三模型微调后的覆盖集成,引入16次修正(+2.4pp),由结构化多智能体系统(构建者、批评者、回归守护者)选择,单轮增益超过此前8次尝试总和。
原文摘要 · Abstract (English)
This paper describes our system for classifying psychological defense mechanisms in emotional support dialogues using the Defense Mechanism Rating Scales (DMRS), placing second (F1 0.406) among 64 teams. A central insight is that defense mechanisms are defined by what is absent: missing affect, blocked cognition, denied reality. We encode this as an affect-cognition integration spectrum in prompt-level clinical rules, which account for the largest single gain (+11.4pp F1). Our architecture is a multi-phase deliberative council of Gemini 2.5 agents where class-specific advocates rate evidence strength rather than voting, achieving F1 0.382 with no fine-tuning - a top-5 result on its own. We find, however, that the council is confidently wrong about minority classes: 59-80% of stable minority predictions are incorrect, driven by a systematic "L7 attractor" in which emotional content defaults to the majority class. A targeted override ensemble from three fine-tuned Qwen3.5 models applies 16 overrides (+2.4pp), selected by a structured multi-agent system (builder, critic, regression guard) that produced a larger F1 gain in one iteration than 8 prior attempts combined.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。