arXiv:2605.19846cs.CVcs.AI2026-05被引 2

构建细粒度动作理解基准,评估视觉语言模型对复杂人类行为的感知能力。

FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding

论文配图:FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
图 1 · 摘自论文原文
  • 设计覆盖64段长视频的密集问答数据集,聚焦人物动作与交互细节。
  • 发现开源模型在多人场景空间推理和细微动作区分上表现显著不足。
  • 提出FineAgent框架,通过定位器与描述器模块提升模型细粒度理解能力。

视觉语言模型(VLMs)在通用视频理解中表现优异,但在需要精细解析人类动作与互动的实际应用中仍面临挑战。现有以人为中心的评测基准多关注公平性、情绪识别等维度,但缺乏长视频、高密度问答与帧级时空定位的结合。为此,我们提出FineBench,一个专为细粒度人类行为理解设计的视频问答(VQA)基准。该数据集包含199,420个多项选择题,覆盖64段每段15分钟的长视频,聚焦人物运动、人际互动及物体操作,包括组合动作。实验表明,尽管闭源模型如GPT-5表现良好,当前开源VLMs在多人场景的空间推理和细微动作差异辨别上明显落后。针对此问题,我们提出FineAgent,一种基于定位器与描述器的模块化增强框架。实验显示,FineAgent可稳定提升多种开源VLM在FineBench上的表现。FineBench为未来细粒度人本视频理解研究提供严格测试平台,FineAgent则为现有VLM推理能力提升提供实用方案。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained comprehension crucial for real-world applications requiring nuanced interpretation of human actions and interactions. While some recent human-centric benchmarks evaluate aspects of model behaviour such as fairness/ethics, emotion perception, and broader human-centric metrics, they do not combine long-form videos, very dense QA coverage, and frame-level spatial/temporal grounding at scale. To bridge this gap, we introduce FineBench, a human-centric video question answering (VQA) benchmark specifically designed to assess fine-grained understanding. FineBench comprises 199,420 multiple-choice QA pairs densely annotated across 64 long-form videos (15 minutes each), focusing on detailed person movement, person interaction, and object manipulation, including compositional actions. Our extensive evaluation reveals that while proprietary models like GPT-5 achieve respectable performance, current open-source VLMs significantly underperform, struggling particularly with spatial reasoning in multi-person scenes and distinguishing subtle differences in human movements and interactions. To address these identified weaknesses, we propose FineAgent, a modular framework that enhances VLMs by leveraging a Localizer and a Descriptor. Experiments show that FineAgent consistently improves the performance of various open VLMs on FineBench. FineBench provides a rigorous testbed for future research into fine-grained human-centric video understanding, while FineAgent offers a practical approach to enhance such reasoning in current VLMs. Project page and code at https://joslefaure.github.io/assets/html/finebench.html.

视频理解细粒度分析多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。