用更细粒度的人格层次提升语音视频交互中的性格识别效果
Enhancing Personality Recognition by Comparing the Predictive Power of Traits, Facets, and Nuances
- 采用大五人格的细分层级(特质、维度、细微特征)作为标签
- 在音视频数据上,细微特征模型误差降低74%
- 适合做性格分析与跨模态行为建模的研究者参考
人格是复杂的分层结构,通常通过项目问卷汇总成宽泛特质得分进行评估。人格识别模型旨在从多种行为数据中推断人格特质。然而,依赖宽泛特质得分作为真实标签,加之训练数据有限,导致模型泛化能力差,因为相似特质得分可能由不同情境下的多样行为表现。本文探索大五人格模型中更细粒度的层级——维度和细微特征对预测性能的影响,以提升从音视频交互数据中识别人格的效果。基于UDIVA v0.5数据集,我们训练了一个包含跨模态(音视频)和跨主体(配对感知)注意力机制的Transformer模型。结果表明,细微特征级别的模型始终优于维度和特质级别模型,在不同交互场景下均将均方误差降低最多达74%。
原文摘要 · Abstract (English)
Personality is a complex, hierarchical construct typically assessed through item-level questionnaires aggregated into broad trait scores. Personality recognition models aim to infer personality traits from different sources of behavioral data. However, reliance on broad trait scores as ground truth, combined with limited training data, poses challenges for generalization, as similar trait scores can manifest through diverse, context dependent behaviors. In this work, we explore the predictive impact of the more granular hierarchical levels of the Big-Five Personality Model, facets and nuances, to enhance personality recognition from audiovisual interaction data. Using the UDIVA v0.5 dataset, we trained a transformer-based model including cross-modal (audiovisual) and cross-subject (dyad-aware) attention mechanisms. Results show that nuance-level models consistently outperform facet and trait-level models, reducing mean squared error by up to 74% across interaction scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。