用多模态大模型提升第一视角视频问答,效果超越现有基准。
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
- 评估四个主流多模态大模型在第一视角视频上的表现。
- 微调后的Video-LLaVa和Qwen2-VL在开集问答上提升2.6%,闭集问答提升13%。
- 发现模型在空间推理和细粒度物体识别上仍存短板,适合研究视觉语言模型者参考。
第一视角视频问答需处理长时序推理、第一人称视角及频繁相机运动等挑战。本文系统评估了四种主流多模态大模型(GPT-4o、Gemini-1.5-Pro、Video-LLaVa-7B 和 Qwen2-VL-7B-Instruct)在改进版 QaEgo4Dv2 数据集上的表现,该数据集源自 QaEgo4D 并降低了标注噪声,支持更可靠的对比。采用零样本与微调两种方法,在 OpenQA 与 CloseQA 设置下进行测试。结果表明,微调后的 Video-LLaVa-7B 与 Qwen2-VL-7B-Instruct 达到新最优性能,开集问答提升最高达 +2.6%(ROUGE/METEOR),闭集问答提升 +13% 准确率。同时通过详尽错误分析发现,模型在空间推理与细粒度物体识别方面仍存在显著不足,是未来优化的关键方向。
原文摘要 · Abstract (English)
Egocentric Video Question Answering (QA) requires models to handle long-horizon temporal reasoning, first-person perspectives, and specialized challenges like frequent camera movement. This paper systematically evaluates both proprietary and open-source Multimodal Large Language Models (MLLMs) on QaEgo4Dv2 - a refined dataset of egocentric videos derived from QaEgo4D. Four popular MLLMs (GPT-4o, Gemini-1.5-Pro, Video-LLaVa-7B and Qwen2-VL-7B-Instruct) are assessed using zero-shot and fine-tuned approaches for both OpenQA and CloseQA settings. We introduce QaEgo4Dv2 to mitigate annotation noise in QaEgo4D, enabling more reliable comparison. Our results show that fine-tuned Video-LLaVa-7B and Qwen2-VL-7B-Instruct achieve new state-of-the-art performance, surpassing previous benchmarks by up to +2.6% ROUGE/METEOR (for OpenQA) and +13% accuracy (for CloseQA). We also present a thorough error analysis, indicating the model's difficulty in spatial reasoning and fine-grained object recognition - key areas for future improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。