用协作提示缓解大模型视频问答中的遗忘问题
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
- 设计三类提示协同捕捉问题上下文、视觉内容与时间动态
- 在NExT-QA和DramaQA上分别达到55.14%和71.24%准确率
- 适合需要持续学习新视频任务的智能问答系统
近年来,线上视频内容激增凸显了静态视频问答(VideoQA)模型在固定数据集上训练的局限性,难以适应新内容提出的新问题或新任务。本文探索在持续学习框架下的视频问答挑战,实证发现:对大语言模型(LLM)逐序列微调常导致灾难性遗忘。为此,我们提出协作提示(ColPro),融合特定问题约束提示、知识获取提示与视觉时序感知提示,旨在捕捉视频问答中的文本问题上下文、视觉内容及视频时序动态,这一视角在以往研究中未被充分关注。在NExT-QA与DramaQA数据集上的实验结果表明,ColPro优于现有方法,在NExT-QA上达到55.14%准确率,在DramaQA上达到71.24%准确率,验证了其实际应用价值与有效性。
原文摘要 · Abstract (English)
In recent years, the rapid increase in online video content has underscored the limitations of static Video Question Answering (VideoQA) models trained on fixed datasets, as they struggle to adapt to new questions or tasks posed by newly available content. In this paper, we explore the novel challenge of VideoQA within a continual learning framework, and empirically identify a critical issue: fine-tuning a large language model (LLM) for a sequence of tasks often results in catastrophic forgetting. To address this, we propose Collaborative Prompting (ColPro), which integrates specific question constraint prompting, knowledge acquisition prompting, and visual temporal awareness prompting. These prompts aim to capture textual question context, visual content, and video temporal dynamics in VideoQA, a perspective underexplored in prior research. Experimental results on the NExT-QA and DramaQA datasets show that ColPro achieves superior performance compared to existing approaches, achieving 55.14\% accuracy on NExT-QA and 71.24\% accuracy on DramaQA, highlighting its practical relevance and effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。