研究大模型识别足球赛关键瞬间的能力,发现其表现接近随机。
Where is the multimodal goal post? On the Ability of Foundation Models to Recognize Contextually Important Moments
- 用比赛精彩集锦隐含偏好构建新数据集
- 主流多模态模型识别关键片段准确率接近随机
- 模型依赖单一模态,跨模态融合能力弱
基础模型广泛应用于从时序多模态事件生成语言的任务中。本文研究模型识别视频中最重要的子事件的能力,这是叙述或总结多模态事件的基本前提。我们聚焦于足球比赛,评估模型区分重要与非重要子事件的能力。为此,我们利用比赛精彩集锦中隐含的人类重要性偏好构建了一个新数据集,无需额外标注成本。基于该数据集,我们对比了多个最先进的多模态模型,结果表明它们的表现仅略高于随机水平。对模型的深入分析显示,它们倾向于依赖单一主导模态,且在整合多源信息方面效果不佳。研究强调了处理多模态数据样本级异质性的模块化架构的重要性,以及需要互补训练方法以最大化跨模态协同效应。
原文摘要 · Abstract (English)
Foundation models are used for many real-world applications involving language generation from temporally-ordered multimodal events. In this work, we study the ability of models to identify the most important sub-events in a video, which is a fundamental prerequisite for narrating or summarizing multimodal events. Specifically, we focus on football games and evaluate models on their ability to distinguish between important and non-important sub-events in a game. To this end, we construct a new dataset by leveraging human preferences for importance implicit in football game highlight reels, without any additional annotation costs. Using our dataset, we compare several state-of-the-art multimodal models and show that they are not far from chance level performance. Analyses of models beyond standard evaluation metrics reveal their tendency to rely on a single dominant modality and their ineffectiveness in synthesizing necessary information from multiple sources. Our findings underline the importance of modular architectures that can handle sample-level heterogeneity in multimodal data and the need for complementary training procedures that can maximize cross-modal synergy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。