arXiv:2608.21022cs.CVcs.MM2026-08中稿 · ACM Multimedia 202…被引 1

无需训练,用提示词让多模态大模型精准理解微动作背后的身心状态。

Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding

论文配图:Recognition-Conditioned Reasoning: A Training-Free Multimodal-LLM Pipeline for Fine-Grained Micro-Action Understanding
图 1 · 摘自论文原文
  • 按任务类型动态分配最优模型:识别用判别式,描述推理用生成式。
  • 在开放任务上得分2.68(满分5),显著优于第二名的1.44。
  • 适合需要零训练、高解释性的微动作分析场景。

微动作是人类无意识做出的细微、短暂且幅度小的身体动作,如手指轻颤或头部微倾,能可靠反映情绪与心理状态。理解这些动作不仅需打标签,还需描述具体动作部位并合理推断其背后原因。我们提出一种完全无需训练、仅依赖提示词的系统,在不允许微调和真实标签监督的MAC 2026微动作挑战赛细粒度理解赛道中夺冠。系统基于冻结的多模态大语言模型(MLLM),将八个子任务动态分配给最适配的模型:判别式MLLM用于封闭式识别,生成式MLLM用于开放式描述与推理。该架构在开放任务上表现显著更优,平均得分达2.68(五分制),远超第二名的1.44。

原文摘要 · Abstract (English)

Micro-actions are subtle, short, low-amplitude body movements, such as a fidgeting hand or a slight head tilt, that humans perform with little conscious intent yet that reliably leak emotional and psychological state. Understanding them goes beyond assigning a label: a model must also describe which body parts move and reason, faithfully, about why a clip warrants a particular fine-grained category. We present the training-free, prompt-only system that won first place in the fine-grained understanding track (MA-Bench) of the MAC~2026 Micro-Action Challenge, where both fine-tuning and ground-truth supervision are disallowed. Built entirely upon frozen multimodal large language models (MLLMs), the system dynamically routes each of the eight sub-tasks to the MLLM empirically best suited for that task: a discriminative MLLM for closed-ended recognition tasks and a generative MLLM for open-ended description and reasoning tasks. This architecture achieves a statistically significant performance advantage on open-ended tasks, attaining an average score of 2.68 (on a five-point scale) compared to 1.44 for the second-best approach.

微动作理解多模态大模型零样本推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。