多模态人格识别模型提升非理想数据下的表现鲁棒性
Multi-modal expressive personality recognition in data non-ideal audiovisual based on multi-scale feature enhancement and modal augment
- 通过跨注意力融合视听特征,实现多尺度特征增强
- 在ChaLearn数据集上达成0.916的平均大五人格识别准确率
- 适合需要抗噪声、缺模态等实际场景的人格分析应用
自动人格识别是计算机科学与心理学交叉领域的研究热点,在人机交互、个性化服务等场景中具有广泛应用。本文构建了一个端到端的视听多模态人格识别网络,通过特征级融合与跨注意力机制有效整合视觉和听觉信息;提出多尺度特征增强模块,强化有效特征表达并抑制冗余信息干扰;训练阶段引入模态增强策略,模拟模态缺失、噪声等非理想数据情况,提升模型对复杂环境的适应能力。实验结果表明,该方法在ChaLearn First Impression数据集上平均大五人格识别准确率达0.916,优于现有音频-视觉双模态方法。消融实验证明各模块及增强策略对性能有显著贡献。推理阶段模拟六种非理想数据场景,验证了模态增强策略对模型鲁棒性的提升效果。
原文摘要 · Abstract (English)
Automatic personality recognition is a research hotspot in the intersection of computer science and psychology, and in human-computer interaction, personalised has a wide range of applications services and other scenarios. In this paper, an end-to-end multimodal performance personality is established for both visual and auditory modal datarecognition network , and the through feature-level fusion , which effectively of the two modalities is carried out the cross-attention mechanismfuses the features of the two modal data; and a is proposed multiscale feature enhancement modalitiesmodule , which enhances for visual and auditory boththe expression of the information of effective the features and suppresses the interference of the redundant information. In addition, during the training process, this paper proposes a modal enhancement training strategy to simulate non-ideal such as modal loss and noise interferencedata situations , which enhances the adaptability ofand the model to non-ideal data scenarios improves the robustness of the model. Experimental results show that the method proposed in this paper is able to achieve an average Big Five personality accuracy of , which outperforms existing 0.916 on the personality analysis dataset ChaLearn First Impressionother methods based on audiovisual and audio-visual both modalities. The ablation experiments also validate our proposed , respectivelythe contribution of module and modality enhancement strategy to the model performance. Finally, we simulate in the inference phase multi-scale feature enhancement six non-ideal data scenarios to verify the modal enhancement strategy's improvement in model robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。