arXiv:2503.12912cs.CVcs.LG2025-03AAAI被引 7

用人体姿态提升性格预测准确率,提出新数据集与心理启发模型

Pose as a Modality: A Psychology-Inspired Network for Personality Recognition with a New Multimodal Dataset

  • 基于心理学理论设计网络,融合姿态、视觉等多模态特征
  • 在287人数据集上性能超越主流模型,姿态信息贡献显著
  • 适合研究性格识别、人机交互与心理计算的学者参考

近年来,从多模态数据预测大五人格特质受到人工智能领域广泛关注。然而,现有计算模型性能仍不理想。心理学研究表明姿态与人格特质强相关,但以往研究多忽略姿态数据。为此,我们构建了一个包含全身姿态的新多模态数据集,收录287名参与者完成36个问题的虚拟面试视频,附有自评大五人格评分作为标签。为有效利用该数据,提出心理学启发网络(PINet),包含三个模块:多模态特征感知(MFA)、多模态特征交互(MFI)和心理学引导的模态相关性损失(PIMC Loss)。MFA模块采用视觉Mamba块捕捉与人格相关的完整视觉特征;MFI模块高效融合多模态信息;PIMC Loss基于心理理论,指导模型对不同人格维度强调不同模态。实验表明,PINet优于多个前沿基线模型,且三个模块对整体性能贡献接近。引入姿态数据显著提升性能,姿态模态在五个模态中居中等重要性。该工作填补了缺乏全身姿态数据的人格研究空白,为提升性格预测模型精度提供新路径,凸显将心理洞察融入AI框架的重要性。

原文摘要 · Abstract (English)

In recent years, predicting Big Five personality traits from multimodal data has received significant attention in artificial intelligence (AI). However, existing computational models often fail to achieve satisfactory performance. Psychological research has shown a strong correlation between pose and personality traits, yet previous research has largely ignored pose data in computational models. To address this gap, we develop a novel multimodal dataset that incorporates full-body pose data. The dataset includes video recordings of 287 participants completing a virtual interview with 36 questions, along with self-reported Big Five personality scores as labels. To effectively utilize this multimodal data, we introduce the Psychology-Inspired Network (PINet), which consists of three key modules: Multimodal Feature Awareness (MFA), Multimodal Feature Interaction (MFI), and Psychology-Informed Modality Correlation Loss (PIMC Loss). The MFA module leverages the Vision Mamba Block to capture comprehensive visual features related to personality, while the MFI module efficiently fuses the multimodal features. The PIMC Loss, grounded in psychological theory, guides the model to emphasize different modalities for different personality dimensions. Experimental results show that the PINet outperforms several state-of-the-art baseline models. Furthermore, the three modules of PINet contribute almost equally to the model's overall performance. Incorporating pose data significantly enhances the model's performance, with the pose modality ranking mid-level in importance among the five modalities. These findings address the existing gap in personality-related datasets that lack full-body pose data and provide a new approach for improving the accuracy of personality prediction models, highlighting the importance of integrating psychological insights into AI frameworks.

性格识别多模态学习姿态分析心理学启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。