用常识和隐喻知识提升模型对心理梗图的识别能力
Figurative-cum-Commonsense Knowledge Infusion for Multimodal Mental Health Meme Classification
- 构建含6类焦虑症状的AxiOM数据集,融合常识与隐喻理解
- 在加权F1上提升4.20%,在抑郁症状识别中验证泛化性
- 适合研究心理健康、多模态理解或隐喻推理的学者
近年来,用户通过非传统方式如表情包表达心理健康问题,常借助其中的隐喻复杂性展现心理困扰。尽管人类能依靠常识理解这些表达,当前多模态语言模型(MLMs)仍难以捕捉表情包中的隐喻特征。为此,我们基于GAD焦虑量表构建了新数据集AxiOM,将表情包细分为六类焦虑症状。提出一种融合常识与领域知识的框架M3H,以增强模型对隐喻语言和常识的理解能力。在6个基线模型(共20种变体)上进行评估,量化与定性结果均显示显著提升,加权F1指标分别提高4.20%和4.66%。进一步在公开数据集RESTORE上测试其在抑郁症状识别中的泛化能力,并开展详尽消融实验,验证各模块贡献。结果揭示现有模型局限,凸显常识知识对隐喻理解的关键作用。
原文摘要 · Abstract (English)
The expression of mental health symptoms through non-traditional means, such as memes, has gained remarkable attention over the past few years, with users often highlighting their mental health struggles through figurative intricacies within memes. While humans rely on commonsense knowledge to interpret these complex expressions, current Multimodal Language Models (MLMs) struggle to capture these figurative aspects inherent in memes. To address this gap, we introduce a novel dataset, AxiOM, derived from the GAD anxiety questionnaire, which categorizes memes into six fine-grained anxiety symptoms. Next, we propose a commonsense and domain-enriched framework, M3H, to enhance MLMs' ability to interpret figurative language and commonsense knowledge. The overarching goal remains to first understand and then classify the mental health symptoms expressed in memes. We benchmark M3H against 6 competitive baselines (with 20 variations), demonstrating improvements in both quantitative and qualitative metrics, including a detailed human evaluation. We observe a clear improvement of 4.20% and 4.66% on weighted-F1 metric. To assess the generalizability, we perform extensive experiments on a public dataset, RESTORE, for depressive symptom identification, presenting an extensive ablation study that highlights the contribution of each module in both datasets. Our findings reveal limitations in existing models and the advantage of employing commonsense to enhance figurative understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。