小模型通过分步思考实现精准视频辅助决策,效果媲美大模型。
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

- 引入分步推理训练数据,让模型先暂停思考再输出动作建议。
- 40亿参数模型在场景理解上达58.0%,仅用大模型1/59的参数。
- 无需额外训练即可泛化到多个新任务,适合实用型智能助手场景。
当前视觉语言模型在视频理解中存在接地推理弱、时间一致性差和上下文规划不足的问题。我们提出 pause-and-think-T,一个以推理为核心的训练数据集,引导模型暂停、基于视觉证据推理,并生成简洁可操作的回应。该数据集促进生成前的结构化推理,使模型更接近人类的场景化辅助行为。我们微调了一个40亿参数的小模型,并在 pause-and-think-B 基准上评估其在上下文理解与目标规划任务的表现。该模型以59倍更少参数(40亿对比2350亿)达到58.0%准确率,超越GPT-4o,在场景理解上与GPT-5.2持平。此外,在EgoThink与TempCompass等分布外数据集上,其在可操作性、辅助判断、归因识别、情境推理与时间顺序理解等方面均有显著提升,且无需针对特定基准训练。结果表明,针对性的推理监督可使小型模型实现行动可执行、视觉感知强的指导能力,并具备良好泛化性,无需扩大模型规模。
原文摘要 · Abstract (English)
Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-and-think-T, a reasoning-centric training dataset that encourages models to pause, reason over visual evidence, and produce concise, actionable responses. The dataset promotes structured reasoning prior to answer generation, guiding models toward human-like, scene-grounded assistance. We fine-tune a compact 4B-parameter model and evaluate it on our pause-and-think-B benchmark targeting contextual understanding and goal planning tasks. The model achieves 58.0% accuracy at 59x fewer parameters than Qwen3-VL-235B (58.9%), matching GPT-5.2 on scene understanding and surpassing GPT-4o. Beyond our benchmark, it also shows strong out-of-distribution performance on EgoThink and TempCompass, with substantial gains in affordance, assistance, attribution recognition, situated reasoning, and temporal order, without benchmark-specific training. Our results indicate that targeted reasoning supervision enables compact models to deliver actionable, visually grounded guidance while generalizing beyond training data, without requiring large-scale model expansion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。