构建细粒度微动作理解基准,推动多模态大模型感知细微行为
MA-Bench: Towards Fine-grained Micro-Action Understanding
- 设计三层次评估架构,系统测试微动作感知、关系理解与解释推理
- 包含12000个问答对,23个模型在细粒度动作识别上表现不佳
- 提供20.5万视频的标注训练集,可显著提升模型微动作理解能力
随着多模态大语言模型(MLLMs)的快速发展,其在微动作理解——人类情绪分析中的关键任务——方面的潜力尚未被充分探索,主要因缺乏专用基准。为此,我们提出MA-Bench,一个包含1,000个视频和三层评估架构的基准,逐步考察微动作感知、关系理解与解释推理。该基准包含12,000个结构化问答对,支持对识别准确率与动作解释能力的系统评估。23个代表性MLLM的实验结果表明,模型在捕捉运动精细度与身体部位动态方面仍存在显著挑战。为此,我们进一步构建了包含20.5万视频的MA-Bench-Train训练数据集,所有视频配有结构化微动作描述。基于该数据集微调的Qwen3-VL-8B模型在微动作推理与解释任务中表现明显提升。本工作旨在为推进MLLMs理解细微微动作与人类行为奠定基础。
原文摘要 · Abstract (English)
With the rapid development of Multimodal Large Language Models (MLLMs), their potential in Micro-Action understanding, a vital role in human emotion analysis, remains unexplored due to the absence of specialized benchmarks. To tackle this issue, we present MA-Bench, a benchmark comprising 1,000 videos and a three-tier evaluation architecture that progressively examines micro-action perception, relational comprehension, and interpretive reasoning. MA-Bench contains 12,000 structured question-answer pairs, enabling systematic assessment of both recognition accuracy and action interpretation. The results of 23 representative MLLMs reveal that there are significant challenges in capturing motion granularity and fine-grained body-part dynamics. To address these challenges, we further construct MA-Bench-Train, a large-scale training corpus with 20.5K videos annotated with structured micro-action captions for fine-tuning MLLMs. The results of Qwen3-VL-8B fine-tuned on MA-Bench-Train show clear performance improvements across micro-action reasoning and explanation tasks. Our work aims to establish a foundation benchmark for advancing MLLMs in understanding subtle micro-action and human-related behaviors. Project Page: https://MA-Bench.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。