构建视频问答基准,测试模型识别相似运动的细微差别能力。
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
- 设计多选题视频问答数据集,聚焦体育领域复杂动作
- 顶尖模型准确率45.52%,低于人类非专家61.64%
- 强调高帧采样对识别细微动作的关键作用
现实世界中存在大量跨领域的复杂动作,同一领域内动作常高度相似,给深度模型识别带来挑战。为评估多模态基础模型在识别此类动作上的表现,我们提出ActionAtlas v1.0,一个基于短视频的多项选择视频问答基准,涵盖56项运动、934个视频、580种独特动作,共包含1896个选项动作。每个视频配一个问题,要求识别特定人物在特定时间点的动作描述。与现有仅覆盖简单动作的基准不同,ActionAtlas聚焦于视觉上相似的精细动作,严格检验模型辨别能力。我们在该基准上评估了开源与专有基础模型,发现最佳模型GPT-4o最高准确率为45.52%;而未受过专业训练的众包人员,在提供动作描述的情况下达到61.64%准确率,随机猜测约为21%。结果表明,高帧采样率对准确识别至关重要,但部分领先专有模型(如Gemini)默认配置中未包含此功能。
原文摘要 · Abstract (English)
Our world is full of varied actions and moves across specialized domains that we, as humans, strive to identify and understand. Within any single domain, actions can often appear quite similar, making it challenging for deep models to distinguish them accurately. To evaluate the effectiveness of multimodal foundation models in helping us recognize such actions, we present ActionAtlas v1.0, a multiple-choice video question answering benchmark featuring short videos across various sports. Each video in the dataset is paired with a question and four or five choices. The question pinpoints specific individuals, asking which choice "best" describes their action within a certain temporal context. Overall, the dataset includes 934 videos showcasing 580 unique actions across 56 sports, with a total of 1896 actions within choices. Unlike most existing video question answering benchmarks that only cover simplistic actions, often identifiable from a single frame, ActionAtlas focuses on intricate movements and rigorously tests the model's capability to discern subtle differences between moves that look similar within each domain. We evaluate open and proprietary foundation models on this benchmark, finding that the best model, GPT-4o, achieves a maximum accuracy of 45.52%. Meanwhile, Non-expert crowd workers, provided with action description for each choice, achieve 61.64% accuracy, where random chance is approximately 21%. Our findings with state-of-the-art models indicate that having a high frame sampling rate is important for accurately recognizing actions in ActionAtlas, a feature that some leading proprietary video models, such as Gemini, do not include in their default configuration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。