用语言查询精准定位人物视频中的精彩帧,提升细粒度理解能力。
ShotVL: Human-Centric Highlight Frame Retrieval via Language Queries
- 基于语言查询与人体动作语义,实现帧级精确定位。
- 在BestShot基准上比InternVL提升52%,在THUMOS14上提升57%。
- 适合需要高精度视频摘要与人机交互的应用场景。
现有针对人物为中心的视频理解研究多聚焦于特定时刻或完整视频分析,但许多应用需达到帧级更高精度。本文提出新任务BestShot,旨在通过语言查询定位人物视频中的精彩帧,要求兼具深层语义理解与精确时间定位能力。为此,我们构建了BestShot基准,结合人工标注的精彩帧、详细文本描述及持续时间标注,描述涵盖视觉内容、细粒度动作和人体姿态三要素。为支持该任务,我们收集两个数据集:(i) ShotGPT4o(由GPT-4o生成),(ii) Image-SMPLText(利用PoseScript和已有姿态估计数据集,实现大规模精准帧级姿态描述)。基于此,我们提出针对BestShot任务微调的ShotVL模型,基于InternVL架构。实验表明,ShotVL在BestShot基准上相比InternVL提升52%,在THUMOS14基准上提升57%,同时保持通用图像分类与检索的SOTA性能。
原文摘要 · Abstract (English)
Existing works on human-centric video understanding typically focus on analyzing specific moment or entire videos. However, many applications require higher precision at the frame level. In this work, we propose a novel task, BestShot, which aims to locate highlight frames within human-centric videos via language queries. This task demands not only a deep semantic comprehension of human actions but also precise temporal localization. To support this task, we introduce the BestShot Benchmark. %The benchmark is meticulously constructed by combining human detection and tracking, potential frame selection based on human judgment, and detailed textual descriptions crafted by human input to ensure precision. The benchmark is meticulously constructed by combining human-annotated highlight frames, detailed textual descriptions and duration labeling. These descriptions encompass three critical elements: (1) Visual content; (2) Fine-grained action; and (3) Human Pose Description. Together, these elements provide the necessary precision to identify the exact highlight frames in videos. To tackle this problem, we have collected two distinct datasets: (i) ShotGPT4o Dataset, which is algorithmically generated by GPT-4o and (ii) Image-SMPLText Dataset, a dataset with large-scale and accurate per-frame pose description leveraging PoseScript and existing pose estimation datasets. Based on these datasets, we present a strong baseline model, ShotVL, fine-tuned from InternVL, specifically for BestShot. We highlight the impressive zero-shot capabilities of our model and offer comparative analyses with existing SOTA models. ShotVL demonstrates a significant 52% improvement over InternVL on the BestShot Benchmark and a notable 57% improvement on the THUMOS14 Benchmark, all while maintaining the SOTA performance in general image classification and retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。