通过情绪线索提升视频面试中人格评估的准确性与可解释性
EMMR: Emotion-Mediated Multimodal Reasoning for Personality Assessment in Asynchronous Video Interviews

- 构建双阶段框架,融合视觉音频中的情绪线索
- 在两个数据集上降低MAE、MSE,提升PCC值
- 适合关注人格评估可解释性的研究者与招聘AI开发者
异步视频面试(AVIs)在人格评估中日益普及。尽管大语言模型(LLMs)在基于转录文本的评估中展现潜力,但纯文本方法可能忽略视觉与音频模态传递的非语言行为线索,而这些线索对人格判断至关重要。尤其情绪相关线索能提供重要的社交与情感证据。为此,我们提出EMMR(情绪中介的多模态推理)框架,用于基于多模态大模型(MLLMs)的AVI人格评估。EMMR从多模态数据中提取情绪线索,并将其作为结构化推理的辅助社会行为证据融入评估过程。在OPVA和AVI-6两个数据集上的实验表明,相比基线方法,EMMR在均方误差(MSE)、平均绝对误差(MAE)和皮尔逊相关系数(PCC)上均有提升。进一步分析显示,情绪线索的语义描述能增强评估效果,其质量直接影响评估可靠性。结果表明,将情绪线索整合进多模态推理是提升可解释性与性能的可行方向。
原文摘要 · Abstract (English)
Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment. Recent large language models (LLMs) have shown potential for personality assessment from transcribed interview responses. However, text-centered methods may overlook non-verbal behavioral cues conveyed through visual and audio modalities, even though such cues are highly relevant to personality assessment. In particular, emotion-related cues provide important social and affective evidence for understanding candidates' behavior related to personality traits. Thus, we propose EMMR (Emotion-Mediated Multimodal Reasoning), a two-stage framework for MLLMs-based personality assessment for AVIs. EMMR extracts emotion-related cues from multimodal interview data and incorporates them into personality assessment through structured reasoning as auxiliary social and behavioral evidence. Experiments on two AVIs datasets, OPVA and AVI-6, show that EMMR improves MAE, MSE, and PCC compared with baselines. Further analysis indicates that semantic descriptions of emotion cues enhance personality assessment, while their quality affects personality assessment reliability. These results suggest that integrating emotion-related cues into multimodal reasoning is a promising direction for more interpretable MLLMs-based personality assessment in AVIs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。