构建视频语言基准ExAct,评估模型对专业技能的细粒度理解能力
ExAct: A Video-Language Benchmark for Expert Action Analysis
- 设计3521个专家标注的视频问答对,覆盖6大领域11类技能
- 顶尖模型GPT-4o准确率仅44.70%,远低于人类专家的82.02%
- 适合研究视觉语言模型在物理技能理解上的局限与突破
我们提出ExAct,一个用于专家级熟练人体动作理解的新视频-语言基准。该基准包含3521个由专家精心筛选的视频问答对,涵盖体育、自行车维修、烹饪、健康、音乐和舞蹈6个领域的11种身体活动。每个问题需从五个精心设计的选项中选出正确答案,要求对人类技能有细致入微的理解。在ExAct上评估最新状态的视觉语言模型(VLMs),发现其性能与人类专家相比存在显著差距:最佳模型GPT-4o仅达44.70%准确率,远低于训练过的专家所达到的82.02%。我们认为ExAct将有助于开发和评估在各种物理与流程性领域具备精确人类技能理解能力的VLM。数据集与代码已公开于https://texaser.github.io/exact_project_page/
原文摘要 · Abstract (English)
We present ExAct, a new video-language benchmark for expert-level understanding of skilled physical human activities. Our new benchmark contains 3521 expert-curated video question-answer pairs spanning 11 physical activities in 6 domains: Sports, Bike Repair, Cooking, Health, Music, and Dance. ExAct requires the correct answer to be selected from five carefully designed candidate options, thus necessitating a nuanced, fine-grained, expert-level understanding of physical human skills. Evaluating the recent state-of-the-art VLMs on ExAct reveals a substantial performance gap relative to human expert performance. Specifically, the best-performing GPT-4o model achieves only 44.70% accuracy, well below the 82.02% attained by trained human specialists/experts. We believe that ExAct will be beneficial for developing and evaluating VLMs capable of precise understanding of human skills in various physical and procedural domains. Dataset and code are available at https://texaser.github.io/exact_project_page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。