arXiv:2607.16284cs.CVcs.MM2026-07被引 2

MAC 2026 推出细粒度微动作理解新任务,用多模态大模型评估深层语义解析能力。

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

论文配图:MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding
图 1 · 摘自论文原文
  • 引入多模态大模型辅助评估,实现对微动作语义的深度理解
  • 新增细粒度微动作理解任务,突破传统识别与检测局限
  • 面向人机交互与情感计算,适合视频理解研究者参考

微动作(MAs)是社交互动与情感交流中重要的非语言线索,具有持续时间短、运动模式弱、语义差异细微等特点,难以标准化标注、建模与评估。为推动该领域研究,我们每年组织微动作分析大赛(MAC),提供公开基准平台。前两届确立了微动作识别与检测的标准评估框架,发布可公开获取的数据集与协议。本文介绍与ACM Multimedia 2026同期举办的第三届MAC,在‘从识别迈向细粒度理解’主题下,进一步拓展挑战范围,首次引入细粒度微动作理解任务,借助多模态大语言模型评估模型捕捉细微语义线索、深入解析人类微动作的能力。本文总结了数据集、任务设置、评估协议、竞赛结果及顶尖团队解决方案,并探讨微动作分析的未来方向及其在以人为中心的视频理解中的关键作用。

原文摘要 · Abstract (English)

Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communication. However, their short duration, weak motion patterns, and fine-grained semantic differences make them difficult to annotate, model, and evaluate in a standardized manner. To promote academic research on micro-action analysis, we proposed and have annually organized the Micro-Action Analysis Grand Challenge (MAC) as a public benchmark platform for this emerging field. The first two editions of MAC established standardized evaluation settings for micro-action recognition and detection, providing publicly accessible datasets and protocols. Building upon these editions, this paper presents the 3rd MAC, held in conjunction with ACM Multimedia 2026. Under the theme of moving from recognition to fine-grained micro-action understanding, this edition further expands the scope of the challenge beyond conventional recognition and detection. In particular, we introduce a new task named fine-grained micro-action understanding, evaluated with the assistance of multimodal large language models, aiming to assess models' ability to capture fine-grained semantic cues and interpret subtle human micro-actions at a deeper level. We summarize the datasets, task settings, evaluation protocols, competition results, and representative solutions from top-performing teams. Finally, we discuss future directions for micro-action analysis and its broader role in human-centric video understanding.

微动作分析细粒度理解多模态大模型视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。