arXiv:2510.19451cs.CVcs.MM2025-10中稿 · ACM Multimedia 202…被引 1

用多模态大模型分析涂鸦心理,让机器像专家一样解读绘画

Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis

  • 分层解析涂鸦,从单个元素到整体结构逐级分析
  • 引入心理知识库,通过强化学习生成个体心理特征画像
  • 可解释性强,适用于临床心理评估与情绪理解任务

多模态大语言模型(MLLMs)在客观感知任务中表现卓越,但在主观、情感丰富的心理分析领域应用仍较少。本文提出PICK框架,针对临床常用的房屋-树木-人物(HTP)测试,通过分层分析与知识注入实现心理图像理解。首先将含多个对象的涂鸦分解为语义有意义的子图,构建三级层次表示:单对象层、多对象层和整体层。随后在各层级进行针对性分析,提取视觉线索对应的心理或情绪信息。引入HTP知识库并设计强化学习训练的特征提取模块,生成单对象层的心理画像,融合整体风格与动态对象特征(如房屋、树木、人物),关联特定心理状态。最后整合多维度信息,输出符合专家推理的评估结果。实验表明,PICK显著提升MLLM在心理分析中的能力,并可推广至情绪理解任务,形成通用可解释框架。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated exceptional performance across various objective multimodal perception tasks, yet their application to subjective, emotionally nuanced domains, such as psychological analysis, remains largely unexplored. In this paper, we introduce PICK, a multi-step framework designed for Psychoanalytical Image Comprehension through hierarchical analysis and Knowledge injection with MLLMs, specifically focusing on the House-Tree-Person (HTP) Test, a widely used psychological assessment in clinical practice. First, we decompose drawings containing multiple instances into semantically meaningful sub-drawings, constructing a hierarchical representation that captures spatial structure and content across three levels: single-object level, multi-object level, and whole level. Next, we analyze these sub-drawings at each level with a targeted focus, extracting psychological or emotional insights from their visual cues. We also introduce an HTP knowledge base and design a feature extraction module, trained with reinforcement learning, to generate a psychological profile for single-object level analysis. This profile captures both holistic stylistic features and dynamic object-specific features (such as those of the house, tree, or person), correlating them with psychological states. Finally, we integrate these multi-faceted information to produce a well-informed assessment that aligns with expert-level reasoning. Our approach bridges the gap between MLLMs and specialized expert domains, offering a structured and interpretable framework for understanding human mental states through visual expression. Experimental results demonstrate that the proposed PICK significantly enhances the capability of MLLMs in psychological analysis. It is further validated as a general framework through extensions to emotion understanding tasks.

心理分析多模态模型涂鸦理解人机共情

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。