用故事化和有依据的问答数据提升视频理解能力
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
- 将零散问题转化为连贯叙事和视觉证据链,构建更丰富的训练信号
- 在STAR和NExT-QA上分别达到72.5%和80.8%准确率,提升超4.9个百分点
- 适合追求模型泛化性与可解释性的视频理解研究者
VideoQA模型的表现受限于监督信号的碎片化特性,通常由孤立的事实性问答对构成。这种‘事实集合’范式无法捕捉事件背后的叙事与因果结构,导致模型仅能浅层理解视频内容。为此,本文提出一种合成更丰富监督信号的框架,包含两种互补策略:基于问题的改写(QBP),将现有问答对中的多样化提问(如‘什么、如何、为何’)整合为重构视频事件结构的连贯段落;基于问题的描述生成(QBC),为每个问题生成细粒度视觉推理依据,使答案有具体证据支撑。利用强大生成模型,以合成数据在统一的下一个词预测目标下训练VideoQA模型。在STAR和NExT-QA上的大量实验验证了该方法,显著提升准确率,刷新基准结果,如3B模型在STAR上达72.5%(+4.9%),7B模型在NExT-QA上达80.8%。此外,分析显示,QBP与QBC均大幅增强跨数据集泛化能力,其中QBP还使模型收敛速度提升2.5倍以上。结果表明,从孤立事实转向叙事连贯性与有依据的推理,能实现更准确、高效且通用的训练范式。
原文摘要 · Abstract (English)
The performance of Video Question Answering (VideoQA) models is fundamentally constrained by the nature of their supervision, which typically consists of isolated, factual question-answer pairs. This "bag-of-facts" approach fails to capture the underlying narrative and causal structure of events, limiting models to a shallow understanding of video content. To move beyond this paradigm, we introduce a framework to synthesize richer supervisory signals. We propose two complementary strategies: Question-Based Paraphrasing (QBP), which synthesizes the diverse inquiries (what, how, why) from a video's existing set of question-answer pairs into a holistic narrative paragraph that reconstructs the video's event structure; and Question-Based Captioning (QBC), which generates fine-grained visual rationales, grounding the answer to each question in specific, relevant evidence. Leveraging powerful generative models, we use this synthetic data to train VideoQA models under a unified next-token prediction objective. Extensive experiments on STAR and NExT-QA validate our approach, demonstrating significant accuracy gains and establishing new state-of-the-art results, such as improving a 3B model to 72.5\% on STAR (+4.9\%) and a 7B model to 80.8\% on NExT-QA. Beyond accuracy, our analysis reveals that both QBP and QBC substantially enhance cross-dataset generalization, with QBP additionally accelerating model convergence by over 2.5x. These results demonstrate that shifting data synthesis from isolated facts to narrative coherence and grounded rationales yields a more accurate, efficient, and generalizable training paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。