arXiv:2504.11515cs.CVcs.CL2025-04被引 4

融合多模态信息,提升短视频中人格特质预测准确率。

Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment

  • 构建面部图结构与双流网络,结合GCN和CNN捕捉静态表情特征。
  • 引入双向GRU与注意力机制,有效提取动态时序关键帧表征。
  • 支持视频、音频、文本多模态融合,适合人格分析研究者使用。

自动预测人格特质已成为计算机视觉领域的挑战性问题。本文提出一种面向短视频片段的人格分析多模态特征学习框架。视觉方面,构建面部图结构并设计基于几何的双流网络,融合图卷积网络(GCN)与卷积神经网络(CNN),以捕捉静态面部表情;同时采用ResNet18与VGGFace提取帧级全局场景与面部外观特征。为建模动态时序信息,集成双向门控循环单元(BiGRU)与时间注意力模块,提取显著帧表示。为增强鲁棒性,引入VGGish CNN提取音频特征,XLM-Roberta提取文本特征。最后,通过多模态通道注意力机制融合不同模态信息,并采用多层感知机(MLP)回归模型预测人格特质。实验结果表明,所提框架在性能上优于现有先进方法。

原文摘要 · Abstract (English)

Predicting personality traits automatically has become a challenging problem in computer vision. This paper introduces an innovative multimodal feature learning framework for personality analysis in short video clips. For visual processing, we construct a facial graph and design a Geo-based two-stream network incorporating an attention mechanism, leveraging both Graph Convolutional Networks (GCN) and Convolutional Neural Networks (CNN) to capture static facial expressions. Additionally, ResNet18 and VGGFace networks are employed to extract global scene and facial appearance features at the frame level. To capture dynamic temporal information, we integrate a BiGRU with a temporal attention module for extracting salient frame representations. To enhance the model's robustness, we incorporate the VGGish CNN for audio-based features and XLM-Roberta for text-based features. Finally, a multimodal channel attention mechanism is introduced to integrate different modalities, and a Multi-Layer Perceptron (MLP) regression model is used to predict personality traits. Experimental results confirm that our proposed framework surpasses existing state-of-the-art approaches in performance.

人格分析多模态学习视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。