arXiv:2505.12007cs.CV2025-05

用事件模态提升暗光下单眼表情识别效果

Multi-modal Collaborative Optimization and Expansion Network for Event-assisted Single-eye Expression Recognition

  • 融合事件与图像模态,协同优化模型表达
  • 在低光照下准确率显著优于传统方法
  • 适合低光照环境下的表情识别研究

本文提出多模态协同优化与扩展网络(MCO-E Net),利用事件模态应对单眼表情识别中的低光照、过曝及高动态范围挑战。网络包含两项创新设计:基于Mamba的多模态协同优化Mamba(MCO-Mamba)和异构协同扩展专家混合模型(HCE-MoE)。MCO-Mamba通过双模态信息联合优化,促进模态间语义协同与融合,均衡学习并发挥各自优势。HCE-MoE采用动态路由机制,分配结构各异的专家(深度、注意力、焦点),实现互补语义的协同学习,系统整合多种特征提取范式以全面捕捉表情语义。大量实验表明,所提网络在单眼表情识别任务中表现优异,尤其在低光照条件下具有显著优势。

原文摘要 · Abstract (English)

In this paper, we proposed a Multi-modal Collaborative Optimization and Expansion Network (MCO-E Net), to use event modalities to resist challenges such as low light, high exposure, and high dynamic range in single-eye expression recognition tasks. The MCO-E Net introduces two innovative designs: Multi-modal Collaborative Optimization Mamba (MCO-Mamba) and Heterogeneous Collaborative and Expansion Mixture-of-Experts (HCE-MoE). MCO-Mamba, building upon Mamba, leverages dual-modal information to jointly optimize the model, facilitating collaborative interaction and fusion of modal semantics. This approach encourages the model to balance the learning of both modalities and harness their respective strengths. HCE-MoE, on the other hand, employs a dynamic routing mechanism to distribute structurally varied experts (deep, attention, and focal), fostering collaborative learning of complementary semantics. This heterogeneous architecture systematically integrates diverse feature extraction paradigms to comprehensively capture expression semantics. Extensive experiments demonstrate that our proposed network achieves competitive performance in the task of single-eye expression recognition, especially under poor lighting conditions.

表情识别事件视觉多模态低光照

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。