对比三种学习范式在第一人称视频摘要中的表现,发现通用大模型更优。
Comparing Learning Paradigms for Egocentric Video Summarization
- 比较监督、无监督和提示微调三种方法在第一人称视频上的应用效果
- 通用大模型GPT-4o在小规模数据上优于专用模型Shotluck Holmes和TAC-SUM
- 揭示当前方法对第一人称视角适应性不足,适合关注视觉理解与模型泛化的研究者
本研究考察了监督学习、无监督学习和提示微调三种计算机视觉范式在理解第一人称视频数据方面的能力。具体评估了当前最先进的监督模型Shotluck Holmes、无监督模型TAC-SUM以及提示微调的通用大模型GPT-4o在视频摘要任务中的表现。结果表明,现有最先进模型在第一人称视频上的性能低于第三视角视频,凸显该领域仍需进一步发展。值得注意的是,在受限资源下仅使用Ego-Exo4D数据集的小规模子集进行评估时,提示微调的通用模型GPT-4o表现超越了专门设计的模型,反映出现有方法在应对第一人称视角独特挑战时的局限性。本研究旨在提供一个全面的概念验证分析,推动计算机视觉技术在第一人称视频中的应用与发展。
原文摘要 · Abstract (English)
In this study, we investigate various computer vision paradigms - supervised learning, unsupervised learning, and prompt fine-tuning - by assessing their ability to understand and interpret egocentric video data. Specifically, we examine Shotluck Holmes (state-of-the-art supervised learning), TAC-SUM (state-of-the-art unsupervised learning), and GPT-4o (a prompt fine-tuned pre-trained model), evaluating their effectiveness in video summarization. Our results demonstrate that current state-of-the-art models perform less effectively on first-person videos compared to third-person videos, highlighting the need for further advancements in the egocentric video domain. Notably, a prompt fine-tuned general-purpose GPT-4o model outperforms these specialized models, emphasizing the limitations of existing approaches in adapting to the unique challenges of first-person perspectives. Although our evaluation is conducted on a small subset of egocentric videos from the Ego-Exo4D dataset due to resource constraints, the primary objective of this research is to provide a comprehensive proof-of-concept analysis aimed at advancing the application of computer vision techniques to first-person videos. By exploring novel methodologies and evaluating their potential, we aim to contribute to the ongoing development of models capable of effectively processing and interpreting egocentric perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。