arXiv:2604.12777cs.CVcs.AI2026-04中稿 · ICRA

模仿大脑处理情绪的双流机制,提升动态表情识别效果。

Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling

论文配图:Cognition-Inspired Dual-Stream Semantic Enhancement for Vision-Based Dynamic Emotion Modeling
图 1 · 摘自论文原文
  • 构建双流架构:一处理语言提示预激活,一融合语义知识
  • 在多个野外数据集上达到最优性能,超越现有方法
  • 适合研究人机情绪交互与可解释性模型的学者

人类大脑并非孤立处理面部表情来形成情绪感知,而是通过感官输入与语义、上下文知识的动态分层整合。然而,现有基于视觉的动态情绪建模方法常忽视情绪感知与认知理论。为弥合机器与人类情绪感知的差距,我们提出受认知启发的双流语义增强模型(DuSE)。该模型模拟双流认知架构:第一流为层次化时间提示聚类(HTPC),实现认知启动效应,通过将文本语义与面部动态的细粒度时序特征对齐,调节视觉刺激的处理;第二流为潜在语义情绪聚合器(LSEA),计算模拟知识整合过程,类比概念行为理论,融合感官输入并结合学习到的概念知识,体现海马体与默认模式网络在构建连贯情绪体验中的作用。通过显式建模这些神经认知机制,DuSE 提供更符合神经科学原理且鲁棒的动态面部表情识别框架。在多个具有挑战性的野外基准测试中,实验验证了该认知中心方法的有效性,表明模拟大脑的情绪处理策略能取得最先进性能,并提升模型可解释性。

原文摘要 · Abstract (English)

The human brain constructs emotional percepts not by processing facial expressions in isolation, but through a dynamic, hierarchical integration of sensory input with semantic and contextual knowledge. However, existing vision-based dynamic emotion modeling approaches often neglect emotion perception and cognitive theories. To bridge this gap between machine and human emotion perception, we propose cognition-inspired Dual-stream Semantic Enhancement (DuSE). Our model instantiates a dual-stream cognitive architecture. The first stream, a Hierarchical Temporal Prompt Cluster (HTPC), operationalizes the cognitive priming effect. It simulates how linguistic cues pre-sensitize neural pathways, modulating the processing of incoming visual stimuli by aligning textual semantics with fine-grained temporal features of facial dynamics. The second stream, a Latent Semantic Emotion Aggregator (LSEA), computationally models the knowledge integration process, akin to the mechanism described by the Conceptual Act Theory. It aggregates sensory inputs and synthesizes them with learned conceptual knowledge, reflecting the role of the hippocampus and default mode network in constructing a coherent emotional experience. By explicitly modeling these neuro-cognitive mechanisms, DuSE provides a more neurally plausible and robust framework for dynamic facial expression recognition (DFER). Extensive experiments on challenging in-the-wild benchmarks validate our cognition-centric approach, demonstrating that emulating the brain's strategies for emotion processing yields state-of-the-art performance and enhances model interpretability.

情绪识别双流模型认知启发可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。