arXiv:2606.29900cs.CVcs.AI2026-06被引 1

用面部微表情与文本融合,提升视频面试人格识别准确率

LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion

论文配图:LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion
图 1 · 摘自论文原文
  • 将面部动作单元转为语义描述,与回答文本在大模型中融合
  • 在AVI-6数据集上误差更低,与人工评分相关性更强
  • 方法可解释性强,适合招聘与心理学研究场景

异步视频面试(AVI)中的人格识别日益重要。现有方法多依赖大语言模型分析文字回复,但单模态易丢失信息(如面部线索);而使用全脸图像或稀疏帧的多模态方法又会丢弃对人格评估至关重要的细粒度时序动态。为此,我们提出一种基于大语言模型的框架,将面部动作单元(AUs)序列转化为可解释的文本描述,并与受访者文字回应在大模型中进行语义融合。轻量级回归头将融合后的嵌入转换为连续人格得分,同时保持语义空间完整性。在AVI-6基准上的实验表明,该方法显著优于多数基线,预测误差更低,且在多个性格维度上与人工评分相关性更强。进一步分析显示,由AU生成的语义表示能提供与文字互补的非语言线索。在大模型中解耦语义理解与回归预测也带来更高的训练稳定性和更清晰的可解释性。结果表明,AU-文本融合是一种心理合理且计算高效的视频面试人格识别框架。

原文摘要 · Abstract (English)

Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in modern recruitment. Existing approaches often rely on large language models (LLMs) to analyze textual responses of interviewees in AVI. However, unimodel methods often suffer from information loss (e.g., ignore facial cues). In contrast, multimodal methods that employ full-face images or sparsely sampled frames can discard fine-grained temporal dynamics critical for accurate personality assessment. To overcome these limitations, we propose an LLM-based framework that semantically fuse facial action units (AUs) with textual responses of AVI. AU sequences are first converted into interpretable textual descriptions, which are then fused with participants' textual responses through an LLM. A lightweight regression head transforms the resulting embeddings into continuous personality scores without disrupting the underlying semantic space. Experiments on the AVI-6 benchmark demonstrate consistent improvements over most baselines, with lower prediction errors and stronger correlations with human-rated scores across multiple traits. Further analysis reveals that AU-derived semantic representations offer complementary non-verbal cues to textual responses. Decoupling semantic understanding from regression prediction within the LLM also leads to greater training stability and clearer interpretability. Overall, these findings demonstrate that AU-text fusion provides a psychologically grounded and computationally efficient framework for personality recognition in AVIs.

人格识别多模态融合面部动作单元大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。