用双模式推理提升体育视频理解,无需训练即达顶尖水平
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
- 采用反应式与反思式双重推理机制,适配不同复杂度体育问题
- 在多个体育视频问答数据集上超越现有方法,最高提升12.3%
- 内置跨9项运动的多模态知识图谱,适合体育智能研究者使用
基于大语言模型的视频问答在通用视频理解中展现潜力,但在体育视频这一高度复杂的领域面临挑战。本文提出FineQuest,首个无需训练的框架,借鉴认知科学设计双模式推理:对简单问题采用反应式推理,对复杂问题采用反思式推理。为弥合通用模型与体育领域理解之间的知识差距,FineQuest引入SSGraph——一个覆盖九种体育项目的多模态体育知识场景图,编码视觉实例与领域术语以提升推理精度。此外,我们构建了两个新体育视频问答基准Gym-QA和Diving-QA,源自FineGym与FineDiving数据集,支持多样化评估。FineQuest在这些新基准及现有SPORTU数据集上均达到领先性能,同时保持强大的通用视频问答能力。
原文摘要 · Abstract (English)
Video Question Answering (VideoQA) based on Large Language Models (LLMs) has shown potential in general video understanding but faces significant challenges when applied to the inherently complex domain of sports videos. In this work, we propose FineQuest, the first training-free framework that leverages dual-mode reasoning inspired by cognitive science: i) Reactive Reasoning for straightforward sports queries and ii) Deliberative Reasoning for more complex ones. To bridge the knowledge gap between general-purpose models and domain-specific sports understanding, FineQuest incorporates SSGraph, a multimodal sports knowledge scene graph spanning nine sports, which encodes both visual instances and domain-specific terminology to enhance reasoning accuracy. Furthermore, we introduce two new sports VideoQA benchmarks, Gym-QA and Diving-QA, derived from the FineGym and FineDiving datasets, enabling diverse and comprehensive evaluation. FineQuest achieves state-of-the-art performance on these benchmarks as well as the existing SPORTU dataset, while maintains strong general VideoQA capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。