通过解耦面部动态特征,提升微表情识别的准确性与可解释性。
DEFT-LLM: Disentangled Expert Feature Tuning for Micro-Expression Recognition
- 设计多专家架构,分离结构、动态纹理和运动语义三类特征。
- 在多个基准上达到领先性能,尤其擅长捕捉局部面部动作细节。
- 适用于需要高精度与可解释性的微表情分析场景。
微表情识别对推断真实情绪至关重要。将多模态大语言模型应用于该任务,可实现面部运动的时空分析并提供可解释的描述。然而仍存在两大核心挑战:(1) 静态外观与动态运动线索纠缠,导致模型难以聚焦细微运动;(2) 现有微表情数据集中的文本标签与底层面部肌肉运动对应不全,造成语义鸿沟。为此,我们提出DEFT-LLM,通过多专家解耦实现运动语义对齐。首先构建Uni-MER——一个以运动驱动的指令数据集,利用光流与动作单元(AU)标签双重约束,确保文本与局部面部运动在时空上一致且合理对应。随后设计包含三个专家的架构,将面部动态解耦为独立可解释的表示(结构、动态纹理、运动语义)。通过将Uni-MER中对齐的知识注入DEFT-LLM,方法有效引入物理先验,同时发挥大语言模型的跨模态推理能力,从而精准捕捉细微情绪线索。在多个具有挑战性的微表情识别基准上实验表明,本方法达到当前最优性能,尤其在局部面部运动可解释建模方面表现突出。
原文摘要 · Abstract (English)
Micro expression recognition (MER) is crucial for inferring genuine emotion. Applying a multimodal large language model (MLLM) to this task enables spatio-temporal analysis of facial motion and provides interpretable descriptions. However, there are still two core challenges: (1) The entanglement of static appearance and dynamic motion cues prevents the model from focusing on subtle motion; (2) Textual labels in existing MER datasets do not fully correspond to underlying facial muscle movements, creating a semantic gap between text supervision and physical motion. To address these issues, we propose DEFT-LLM, which achieves motion semantic alignment by multi-expert disentanglement. We first introduce Uni-MER, a motion-driven instruction dataset designed to align text with local facial motion. Its construction leverages dual constraints from optical flow and Action Unit (AU) labels to ensure spatio-temporal consistency and reasonable correspondence to the movements. We then design an architecture with three experts to decouple facial dynamics into independent and interpretable representations (structure, dynamic textures, and motion-semantics). By integrating the instruction-aligned knowledge from Uni-MER into DEFT-LLM, our method injects effective physical priors for micro expressions while also leveraging the cross modal reasoning ability of large language models, thus enabling precise capture of subtle emotional cues. Experiments on multiple challenging MER benchmarks demonstrate state-of-the-art performance, as well as a particular advantage in interpretable modeling of local facial motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。