arXiv:2607.10860cs.CV2026-07

用面部动作单元生成7.5万条合成微表情视频,解决数据稀缺问题。

AU-Guided Synthetic Video Generation for Micro-Expression Recognition

论文配图:AU-Guided Synthetic Video Generation for Micro-Expression Recognition
图 1 · 摘自论文原文
  • 基于面部动作单元设计可控视频生成流程
  • 生成7.5万条跨五类情绪的微表情视频,含元数据与质量评估
  • 训练模型在真实数据集上表现稳定,适合微表情识别研究

微表情识别受限于现有数据集规模小、人群覆盖窄、情感标签有限。本文提出EquiME,一个基于面部动作单元(AU)引导的图像到视频生成构建的合成微表情数据集。EquiME包含从1.5万张源人脸图像生成的7.5万条视频,涵盖五种目标情绪,并附带自动推断的年龄、性别等人口统计学元数据及视频质量度量。通过帧对相似性、空间变化和无参考感知质量指标评估,并在SAMM与CASME II数据集上进行跨数据集微表情识别实验。在EquiME上训练的模型在SAMM和CASME II上均取得有竞争力的性能,且在四种架构间表现波动较小。本文聚焦数据集设计、结构化AU条件生成流程及评估其作为合成微表情识别资源的实证证据。

原文摘要 · Abstract (English)

Micro-expression recognition is limited by the small scale, narrow demographic coverage, and restricted emotion labels of existing datasets. We introduce EquiME, a synthetic micro-expression dataset built from AU-guided image-to-video generation. EquiME contains 75K videos generated from 15K source face images across five target emotions, together with automatically inferred demographic metadata and video-quality measurements. We evaluate EquiME using frame-pair similarity, spatial variation, and no-reference perceptual-quality metrics, together with cross-dataset MER experiments on SAMM and CASME II. Models trained on EquiME achieve competitive cross-dataset performance on SAMM and CASME II and show comparatively low variation across the four evaluated architectures. This paper focuses on the dataset design, the structured AU-conditioning pipeline used for video generation, and the empirical evidence needed to assess EquiME as a synthetic MER resource. Project page: https://kirito-blade.github.io/me-vlm/

微表情识别合成数据视频生成AU建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。