arXiv:2603.08927cs.CVcs.MM2026-03被引 1

挑战微表情理解的视觉问答任务,推动模型在短时与长时视频中识别隐藏情绪。

MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering

  • 设计短时与长时微表情视频问答任务,结合大模型多模态推理能力。
  • 需在真实场景长视频中实现跨时段微表情检测与时间推理。
  • 适合关注情绪识别、多模态大模型应用的研究者参与。

面部微表情(MEs)是人在高压力情境下因情绪波动而出现的自发性面部动作,常被抑制或压抑。近年来,微表情识别、定位与生成已取得显著进展。多模态大语言模型(MLLMs)和大视觉语言模型(LVLMs)凭借强大的多模态推理能力,为提升微表情分析提供了新方向。MEGC2026引入两项任务:(1) 微表情视频问答(ME-VQA),在较短视频序列上通过视觉问答方式探索微表情理解,利用MLLMs或LVLMs处理多种与微表情相关的问题;(2) 微表情长视频问答(ME-LVQA),将VQA扩展至真实场景中的长时间视频序列,要求模型具备时间推理能力并检测长期跨度下的细微表情变化。所有参赛算法需在公开排行榜提交结果。更多信息请访问 https://megc2026.github.io。

原文摘要 · Abstract (English)

Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In recent years, substantial advancements have been made in the areas of ME recognition, spotting, and generation. The emergence of multimodal large language models (MLLMs) and large vision-language models (LVLMs) offers promising new avenues for enhancing ME analysis through their powerful multimodal reasoning capabilities. The ME grand challenge (MEGC) 2026 introduces two tasks that reflect these evolving research directions: (1) ME video question answering (ME-VQA), which explores ME understanding through visual question answering on relatively short video sequences, leveraging MLLMs or LVLMs to address diverse question types related to MEs; and (2) ME long-video question answering (ME-LVQA), which extends VQA to long-duration video sequences in realistic settings, requiring models to handle temporal reasoning and subtle micro-expression detection across extended time periods. All participating algorithms are required to submit their results on a public leaderboard. More details are available at https://megc2026.github.io.

微表情视觉问答多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。