arXiv:2604.11177cs.CV2026-04

探究大模型推理过程如何影响视频理解,发现思考越多未必越好。

Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding

  • 通过分析模型内部推理流,评估其对视频理解的影响
  • 前几百个词的思考带来主要提升,后续增益迅速饱和
  • 轻量版模型更省资源且推理内容更聚焦场景本身

我们评估了内部推理轨迹(称为思维流)对视觉语言模型在视频场景理解中的影响。基于谷歌 Gemini 2.5 Flash 和 Flash Lite 四种配置,在100小时视频中提取的场景上进行实验,提出三个问题:更多思考是否带来更好输出?提升在何处停止?模型实际在想什么?引入三项评估指标:内容性衡量推理流中有多少是有效场景内容而非元评论;思维-最终覆盖率衡量推理内容是否忠实转化为最终输出;主导实体分析识别模型关注的主体、动作和场景。以 GPT-5 作为独立裁判。结果表明,质量提升在前几百个词内即达峰值,后续增益迅速趋缓。Flash Lite 在质量和令牌使用间取得最佳平衡。严格推理预算会导致模型在最终输出中添加未经过推理的内容,形成一种压缩阶段幻觉。尽管属不同模型层级,Flash 与 Flash Lite 产生相似的思维流,但风格差异明显:Flash 偏向讨论自身推理过程,Lite 则专注于描述场景。

原文摘要 · Abstract (English)

We benchmark how internal reasoning traces, which we call thought streams, affect video scene understanding in vision-language models. Using four configurations of Google's Gemini 2.5 Flash and Flash Lite across scenes extracted from 100 hours of video, we ask three questions: does more thinking lead to better outputs, where do the gains stop, and what do these models actually think about? We introduce three evaluation metrics. Contentfulness measures how much of the thought stream is useful scene content versus meta-commentary. Thought-Final Coverage measures how faithfully the thought stream translates into the final output. Dominant Entity Analysis identifies which subjects, actions, and settings the model focuses on. GPT-5 serves as an independent judge. We find that quality gains from additional thinking plateau quickly, with most improvement occurring in the first few hundred tokens. Flash Lite offers the best balance between quality and token usage. Tight reasoning budgets cause the model to add content in the final output that it never reasoned about, a form of compression-step hallucination. Despite being different model tiers, Flash and Flash Lite produce similar thought streams, though they differ in style: Flash discusses its reasoning process, while Lite focuses on describing the scene.

视频理解思维流推理评估Gemini

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。