用光流引导大模型理解微表情,提升情绪识别精度。
MELLM: A Flow-Guided Large Language Model for Micro-Expression Understanding
- 结合光流估计与大模型推理,捕捉细微面部动作。
- 在54,611对图像上训练,显著优于现有光流方法。
- 首个专用于微表情理解的大语言模型,适合情绪分析研究者。
微表情(MEs)是揭示隐藏情绪的短暂且低强度面部动作,在情感计算中至关重要。尽管已有方法在离散情绪分类上取得进展,但大多局限于简单分类,缺乏对细微面部动态和情绪线索的全面理解。多模态大语言模型虽具备推理潜力,仍难以感知这些微妙的情绪行为。为此,我们提出面向微表情理解的大型语言模型(MELLM),融合基于光流的细微运动敏感性与大模型的强大推理能力。具体地,提出一种迭代式、基于形变的光流估计算法MEFlowNet,以精准捕捉面部微动作。为训练与评估,构建了包含54,611对起始-峰值图像的MEFlowDataset,涵盖多样身份与细微面部运动。随后设计光流引导的微表情理解范式:利用MEFlowNet提取的光流信号构建MEU-Instruct指令调优数据集,再对MELLM进行微调,使其能将细微运动模式转化为可读描述并生成情绪推断。实验表明,MEFlowNet在面部及微表情光流估计上显著优于现有方法,而MELLM在多个微表情基准测试中达到最佳性能。据我们所知,本工作首次提出专门针对微表情的光流估计算法(MEFlowNet)与专用大语言模型(MELLM)。
原文摘要 · Abstract (English)
Micro-expressions (MEs), brief and low-intensity facial movements revealing concealed emotions, are crucial for affective computing. Despite notable progress in ME recognition, existing methods are largely confined to discrete emotion classification, lacking the capacity for comprehensive ME Understanding (MEU), particularly in interpreting subtle facial dynamics and underlying emotional cues. While Multimodal Large Language Models (MLLMs) offer potential for MEU with their advanced reasoning abilities, they still struggle to perceive such subtle facial affective behaviors. To bridge this gap, we propose a ME Large Language Model (MELLM) that integrates optical flow-based sensitivity to subtle facial motions with the powerful inference ability of LLMs. Specifically, an iterative, warping-based optical-flow estimator, named MEFlowNet, is introduced to precisely capture facial micro-movements. For its training and evaluation, we construct MEFlowDataset, a large-scale optical-flow dataset with 54,611 onset-apex image pairs spanning diverse identities and subtle facial motions. Subsequently, we design a Flow-Guided Micro-Expression Understanding paradigm. Under this framework, the optical flow signals extracted by MEFlowNet are leveraged to build MEU-Instruct, an instruction-tuning dataset for MEU. MELLM is then fine-tuned on MEU-Instruct, enabling it to translate subtle motion patterns into human-readable descriptions and generate corresponding emotional inferences. Experiments demonstrate that MEFlowNet significantly outperforms existing optical flow methods in facial and ME-flow estimation, while MELLM achieves state-of-the-art accuracy and generalization across multiple ME benchmarks. To the best of our knowledge, this work presents two key contributions: MEFlowNet as the first dedicated ME flow estimator, and MELLM as the first LLM tailored for MEU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。