arXiv:2508.08082cs.CV2025-08被引 1

用状态转移建模微表情时序动态,实现精准定位与分类一体化。

ME-TST+: Micro-expression Analysis via Temporal State Transition with ROI Relationship Awareness

  • 采用状态空间模型替代固定窗口分类,支持变长微表情建模。
  • 在3个公开数据集上达到最佳性能,尤其在短时微表情识别中提升显著。
  • 融合多粒度区域关系与慢快框架,适合视频情绪分析研究者使用。

微表情是反映个体内在情绪、偏好与倾向的重要指标,其分析需在长视频序列中定位微表情时段并识别对应情绪类别。现有深度学习方法多采用滑动窗口分类网络,但固定窗口长度与硬分类策略存在明显局限,且通常将定位与识别视为独立任务,忽略二者关联。为此,本文提出基于状态空间模型的ME-TST与ME-TST+架构,以时间状态转移机制取代传统窗口级分类,实现视频级回归,更精确刻画微表情时序动态,并支持不同持续时间的微表情建模。在ME-TST+中,进一步引入多粒度感兴趣区域(ROI)建模与SlowFast Mamba框架,缓解时序任务中的信息丢失问题。同时,提出特征与结果层面的协同策略,利用定位与识别间的内在联系提升整体性能。大量实验表明,所提方法在多个公开数据集上达到当前最优表现。

原文摘要 · Abstract (English)

Micro-expressions (MEs) are regarded as important indicators of an individual's intrinsic emotions, preferences, and tendencies. ME analysis requires spotting of ME intervals within long video sequences and recognition of their corresponding emotional categories. Previous deep learning approaches commonly employ sliding-window classification networks. However, the use of fixed window lengths and hard classification presents notable limitations in practice. Furthermore, these methods typically treat ME spotting and recognition as two separate tasks, overlooking the essential relationship between them. To address these challenges, this paper proposes two state space model-based architectures, namely ME-TST and ME-TST+, which utilize temporal state transition mechanisms to replace conventional window-level classification with video-level regression. This enables a more precise characterization of the temporal dynamics of MEs and supports the modeling of MEs with varying durations. In ME-TST+, we further introduce multi-granularity ROI modeling and the slowfast Mamba framework to alleviate information loss associated with treating ME analysis as a time-series task. Additionally, we propose a synergy strategy for spotting and recognition at both the feature and result levels, leveraging their intrinsic connection to enhance overall analysis performance. Extensive experiments demonstrate that the proposed methods achieve state-of-the-art performance. The codes are available at https://github.com/zizheng-guo/ME-TST.

微表情分析状态空间模型时序建模视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。