检验强制思维链对视频问答是否真有效,发现链路依赖视频但未提升准确率。
Chains That See, Answers That Don't: A Multi-Aspect Evaluation Recipe for Forced Chain-of-Thought on Video-MME
- 设计三重测试:对比直接回答、思维链、先答后链及无视频条件下的表现
- 思维链虽受视频强影响,但未提升多选题准确率,7B模型甚至略有下降
- 提供可复现的评测脚本和原始输出,适合关注模型可信度的研究者
强制思维链(CoT)常被认为能提升视觉语言模型在视频问答中的可靠性。本文提出一个包含三项诊断的轻量级评估方案:在直接回答、思维链、先答后链及无视频条件下比较准确率;通过反事实视频替换测试思维链的连贯性;使用四层视觉退化阶梯评估鲁棒性。所有测试均采用严格与宽松正则匹配器,并对主族结果进行多重检验校正。在Qwen2.5-VL模型与Video-MME子集上的实验显示,思维链强烈依赖输入视频——视频替换导致链路重叠消失且多数答案翻转,与“模板链”假设相反;但在相同数据上,强制思维链并未提升多项选择题准确率,7B模型在事后主评分器下甚至出现统计显著的小幅下降。研究不主张结论普适于其他模型或数据集,原始响应和单次重计算脚本将随补充材料发布,确保每项数据可追溯复现。
原文摘要 · Abstract (English)
Forced chain-of-thought (CoT) is widely assumed to make vision-language models more reliable on video question answering. We propose a small three-probe evaluation recipe to test that assumption: paired accuracy across direct, CoT, answer-first, and no-video conditions; a counterfactual video-swap diagnostic over the CoT chains; and a four-rung visual-degradation ladder. Each probe is reported under both a strict and a permissive regex scorer, with multiplicity correction over a manuscript-declared primary family. Applied to Qwen2.5-VL on Video-MME subsets, the recipe returns a two-part finding. The CoT chains are strongly video-conditioned: swapping the input video collapses chain overlap and flips most final letters, the opposite of what a "boilerplate-chain" null would predict. Yet on the same data, forced CoT does not improve MCQ accuracy, and on the smaller 7B model it produces a small but statistically supported drop under a post-hoc primary scorer choice. We do not claim this generalizes beyond the Qwen2.5-VL / Video-MME instantiation; the raw responses and a single recomputation script will be released with the supplementary material so every number can be re-derived.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。