用儿童教育视频的结构化内容训练视觉模型,提升空间推理能力。
Structure Over Scale: Learning Visual Reasoning from Pedagogical Video
- 利用儿童节目中的提问-暂停-回答循环结构自动构建1万条带时间对齐的问答数据
- 仅用78小时数据训练,使模型在多个基准上提升超10个百分点
- 适合研究多模态推理、低资源训练与教育类AI的开发者
当前顶尖视觉语言模型在视频基准上表现优异,但在涉及空间关系、导航和物体选择等基础视觉推理任务中表现不佳,甚至不如学龄前儿童。我们提出,儿童教育类视频中嵌入的“情境-提问-暂停-回答”教学结构,提供了天然对齐的推理轨迹:时序同步的视觉线索、问题与答案,这些仅源于刻意的教学设计,无法通过大规模人工标注重建。为此,我们构建了SoSVQA(Structure over Scale Visual Question Answering)基准,从《朵拉探险记》(DoraVQA)和《米老鼠俱乐部屋》(ClubHVQA)中自动提取1万组带精确时间戳对齐的问答对。通过使用组相对策略优化(GRPO)微调Qwen2-VL和Qwen3-VL,利用教育内容中明确的正确性信号与结构化推理路径。尽管仅用78小时儿童电视内容(1万条问答),远少于GPT与Gemini的训练规模,该方法仍显著提升模型泛化能力,在NExT-QA(+19.7)、Video-MME(+10.6)和MotionBench(+4.9)上取得持续改进,性能媲美领先专有系统,证明内容结构可弥补数据规模不足。
原文摘要 · Abstract (English)
State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and object selection that a preschooler solves easily. We hypothesize that the explicit pedagogical structure, specifically the context-question-pause-answer cycles embedded in children's educational video, provides naturally co-aligned reasoning traces: temporally synchronized visual cues, questions, and answers that emerge only from deliberate pedagogical authoring and cannot be practically reconstructed through manual annotation at scale. To test this, we introduce SoSVQA (Structure over Scale Visual Question Answering), a unified benchmark of 10K question-answer pairs automatically extracted from Dora the Explorer (DoraVQA) and Mickey Mouse Clubhouse (ClubHVQA) with precise timestamp alignment, and fine-tune Qwen2-VL and Qwen3-VL using Group Relative Policy Optimization (GRPO) to leverage the clear correctness signals and structured reasoning traces inherent in educational content. Despite training on just 10K QA pairs from 78 hours of children's television, orders of magnitude less data than GPT and Gemini, our approach delivers generalizable performance gains for Qwen-based VLMs, yielding consistent improvements on NExT-QA (+19.7), Video-MME (+10.6), and MotionBench (+4.9), matching the performance of leading proprietary systems and demonstrating that content structure can compensate for content scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。