arXiv:2507.21778cs.CV2025-07被引 9

用大模型融合视觉特征,提升微表情动作单元检测精度。

AU-LLM: Micro-Expression Action Unit Detection via Enhanced LLM-Based Feature Fusion

  • 用多层感知机融合局部纹理与全局语义特征,生成紧凑表示
  • 在CASME II和SAMM数据集上达到新最好效果,跨被试测试也稳定
  • 首次将大模型用于微表情动作单元检测,适合情感计算研究者

微表情动作单元(AUs)的检测是情感计算中的重大挑战,对解析细微、无意识的人类情绪至关重要。尽管大语言模型(LLMs)具备强大的推理能力,但其在细粒度、低强度的微表情AU检测领域尚未被探索。本文首次提出AU-LLM框架,利用LLM检测微表情数据集中细微强度的AUs,解决数据稀缺问题。针对视觉-语言语义鸿沟,设计了增强融合投影器(EFP),通过多层感知机(MLP)将专用3D-CNN主干网络提取的中层(局部纹理)与高层(全局语义)视觉特征融合为单一信息密集型标记。该紧凑表征有效赋能LLM对细微面部肌肉运动进行细致推理。在基准数据集CASME II和SAMM上进行广泛评估,包括严格的留一被试(LOSO)和跨域协议,结果表明AU-LLM达到新的最先进水平,验证了基于大模型推理在微表情分析中的显著潜力与鲁棒性。代码已开源:https://github.com/ZS-liu-JLU/AU-LLMs。

原文摘要 · Abstract (English)

The detection of micro-expression Action Units (AUs) is a formidable challenge in affective computing, pivotal for decoding subtle, involuntary human emotions. While Large Language Models (LLMs) demonstrate profound reasoning abilities, their application to the fine-grained, low-intensity domain of micro-expression AU detection remains unexplored. This paper pioneers this direction by introducing \textbf{AU-LLM}, a novel framework that for the first time uses LLM to detect AUs in micro-expression datasets with subtle intensities and the scarcity of data. We specifically address the critical vision-language semantic gap, the \textbf{Enhanced Fusion Projector (EFP)}. The EFP employs a Multi-Layer Perceptron (MLP) to intelligently fuse mid-level (local texture) and high-level (global semantics) visual features from a specialized 3D-CNN backbone into a single, information-dense token. This compact representation effectively empowers the LLM to perform nuanced reasoning over subtle facial muscle movements.Through extensive evaluations on the benchmark CASME II and SAMM datasets, including stringent Leave-One-Subject-Out (LOSO) and cross-domain protocols, AU-LLM establishes a new state-of-the-art, validating the significant potential and robustness of LLM-based reasoning for micro-expression analysis. The codes are available at https://github.com/ZS-liu-JLU/AU-LLMs.

微表情大模型动作单元视觉融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。