arXiv:2509.14592cs.MMcs.SD2025-09

首个融合音视频的微表情数据集,提升真实场景下情绪识别准确率。

MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion

  • 用非对称交叉注意力融合全局视觉与动态音频信号
  • 在跨被试验证中,音频显著提升微表情识别效果
  • 适合情感计算、人机交互领域研究者参考

微表情是隐藏情绪的重要泄露线索,但现有研究多依赖无声的纯视觉数据。为此,本文提出首个在高压力生态情境中采集的音视频同步微表情数据集MMED,首次记录了微表情伴随的自发语音线索。同时,提出异构多模态融合网络AMF-Net,通过非对称交叉注意力机制,将全局视觉表征与动态音频序列进行有效融合。通过严格的留一被试交叉验证(LOSO-CV)实验,结果明确表明音频提供关键解歧信息,显著增强微表情分析能力。该数据集与方法为微表情识别提供了可靠资源与验证框架。

原文摘要 · Abstract (English)

Micro-expressions (MEs) are crucial leakages of concealed emotion, yet their study has been constrained by a reliance on silent, visual-only data. To solve this issue, we introduce two principal contributions. First, MMED, to our knowledge, is the first dataset capturing the spontaneous vocal cues that co-occur with MEs in ecologically valid, high-stakes interactions. Second, the Asymmetric Multimodal Fusion Network (AMF-Net) is a novel method that effectively fuses a global visual summary with a dynamic audio sequence via an asymmetric cross-attention framework. Rigorous Leave-One-Subject-Out Cross-Validation (LOSO-CV) experiments validate our approach, providing conclusive evidence that audio offers critical, disambiguating information for ME analysis. Collectively, the MMED dataset and our AMF-Net method provide valuable resources and a validated analytical approach for micro-expression recognition.

微表情音视频融合多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。