挑战微表情检测与识别一体化,探索大模型问答能力
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering
- 将微表情定位与识别合并为统一流水线,提升长视频分析效率
- 引入视觉问答任务,利用大模型理解复杂微表情场景
- 面向真实场景下的情绪分析,适合多模态认知研究者
面部微表情(MEs)是人在高压力情境下因情绪波动而产生、却试图抑制的自发性面部动作。近年来,微表情识别、定位与生成领域取得显著进展。然而,传统将定位与识别分开处理的方法在真实长视频场景中表现不佳。与此同时,多模态大语言模型(MLLMs)和大视觉-语言模型(LVLMs)凭借强大的跨模态推理能力,为微表情分析提供了新路径。MEGC 2025 设立两项新任务:(1) 微表情定位-识别(ME-STR),将定位与后续识别整合为统一序列流程;(2) 微表情视觉问答(ME-VQA),通过视觉问答形式,借助 MLLMs 或 LVLMs 回答与微表情相关的多样化问题。所有参赛算法需在测试集上运行并提交结果至排行榜。更多详情见 https://megc2025.github.io。
原文摘要 · Abstract (English)
Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In recent years, substantial advancements have been made in the areas of ME recognition, spotting, and generation. However, conventional approaches that treat spotting and recognition as separate tasks are suboptimal, particularly for analyzing long-duration videos in realistic settings. Concurrently, the emergence of multimodal large language models (MLLMs) and large vision-language models (LVLMs) offers promising new avenues for enhancing ME analysis through their powerful multimodal reasoning capabilities. The ME grand challenge (MEGC) 2025 introduces two tasks that reflect these evolving research directions: (1) ME spot-then-recognize (ME-STR), which integrates ME spotting and subsequent recognition in a unified sequential pipeline; and (2) ME visual question answering (ME-VQA), which explores ME understanding through visual question answering, leveraging MLLMs or LVLMs to address diverse question types related to MEs. All participating algorithms are required to run on this test set and submit their results on a leaderboard. More details are available at https://megc2025.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。