arXiv:2512.04728cs.CVcs.AI2025-12

破解对话中隐性心理状态分析难题,提升视觉情感识别准确率。

Measuring the Unspoken: A Disentanglement Model and Benchmark for Psychological Analysis in the Wild

  • 设计分层视觉编码器MIND,通过动态抑制模糊唇部特征实现视觉解耦。
  • 在新构建的数据集上,微表情检测准确率比现有最优模型提升86.95%。
  • 提供可量化评估框架PRISM,适合研究心理视觉推理的学者使用。

自然场景对话中的生成式心理分析面临两大挑战:(1) 现有视觉语言模型无法解决发音与情绪表达之间的歧义问题,即视觉言语模式会模仿情感表现;(2) 缺乏可验证的评估指标来衡量视觉定位和推理深度。为此,我们提出一个完整生态系统:首先,引入多层级洞察网络(MIND),一种新型分层视觉编码器,通过状态判断模块基于时序特征方差算法抑制模糊的唇部特征,实现显式的视觉解耦。其次,构建了包含专家标注的大型数据集ConvoInsight-DB,涵盖微表情与深层心理推断。第三,设计了心理推理洞察评分度量(PRISM),一种基于专家引导大模型的自动化维度评估框架,用于衡量大型心理视觉模型的多维性能。在我们的PRISM基准测试中,MIND显著优于所有基线,在微表情检测任务上相比前人最佳模型提升86.95%。消融实验确认,状态判断解耦模块是性能跃升的关键因素。代码已开源。

原文摘要 · Abstract (English)

Generative psychological analysis of in-the-wild conversations faces two fundamental challenges: (1) existing Vision-Language Models (VLMs) fail to resolve Articulatory-Affective Ambiguity, where visual patterns of speech mimic emotional expressions; and (2) progress is stifled by a lack of verifiable evaluation metrics capable of assessing visual grounding and reasoning depth. We propose a complete ecosystem to address these twin challenges. First, we introduce Multilevel Insight Network for Disentanglement(MIND), a novel hierarchical visual encoder that introduces a Status Judgment module to algorithmically suppress ambiguous lip features based on their temporal feature variance, achieving explicit visual disentanglement. Second, we construct ConvoInsight-DB, a new large-scale dataset with expert annotations for micro-expressions and deep psychological inference. Third, Third, we designed the Mental Reasoning Insight Rating Metric (PRISM), an automated dimensional framework that uses expert-guided LLM to measure the multidimensional performance of large mental vision models. On our PRISM benchmark, MIND significantly outperforms all baselines, achieving a +86.95% gain in micro-expression detection over prior SOTA. Ablation studies confirm that our Status Judgment disentanglement module is the most critical component for this performance leap. Our code has been opened.

心理分析视觉解耦微表情检测评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。