arXiv:2607.01469cs.CV2026-07

重复实验发现,智能视频问答系统性能排序可能反转,质疑单一评估的可靠性。

Accuracy and Cost Claims Do Not Survive Re-Execution in Agentic VideoQA

论文配图:Accuracy and Cost Claims Do Not Survive Re-Execution in Agentic VideoQA
图 1 · 摘自论文原文
  • 通过重复对比相同任务中的动态与静态推理模型,验证结果可逆性。
  • 两次实验中,动态模型分别领先7.33点和落后4.05点,效果完全反转。
  • 提出REPAIR协议,用于检测评估结果是否具有方向可复现性。

智能视频问答系统通过自适应推理和工具使用生成答案,但当前评估仅进行一次,仅在问题间估计不确定性。这留下一个基本问题:测量的方法效应在重复评估下是否仍成立?我们以Static-SAGE和Dynamic-SAGE为受控案例,对相同的SAGE-Bench问答-视频对进行两次配对比较,保持配置、工具库和评分协议一致。第一次执行中,Dynamic-SAGE比Static-SAGE高出7.33个百分点;第二次则反向为-4.05。两者均在配对分析下显著,支持相反结论。两次执行间的效应变化极显著,远超单次执行内的不确定性。该反转在问题格式、模态、难度和视频时长上均一致,且两评估组均发生显著移动。为此,我们提出REPAIR(重复配对推断可靠性)协议,通过重复配对比较并检验方法效应是否随执行变化,区分方向可复现性与效应大小稳定性。应用于准确率和执行指标,REPAIR揭示三类行为:方向反转、幅度变化和效应衰减;并表明推理轮次和可见工具调用减少,并不意味着原始计算或延迟的可复现降低。执行层面的变动与近期智能视频问答系统的平均提升相当,甚至更大,凸显其重要性。单次执行中的显著性不足以证明报告的方法效应具有可复现性。

原文摘要 · Abstract (English)

Agentic Video Question Answering (VideoQA) systems produce answers through adaptive reasoning and tool-use trajectories, yet standard practice evaluates each system once and estimates uncertainty only across questions. This leaves a basic question untested: would the measured method effect survive if the evaluation were run again? We show that it need not. Using Static-SAGE and Dynamic-SAGE as a controlled case study, we repeat the paired comparison twice on identical SAGE-Bench question-video pairs, holding configuration, tool library, and scoring protocol fixed. In the first execution, Dynamic-SAGE outperforms Static-SAGE by +7.33 accuracy points; in the second, the effect reverses to -4.05. Both are individually significant under paired analysis, supporting opposite conclusions. The change in the paired effect between executions is highly significant and far larger than within-execution uncertainty. The reversal is consistent across question format, modality, difficulty, and video duration, and both evaluation arms move significantly. Motivated by this failure, we introduce REPAIR (REpeated PAired Inference Reliability), a protocol that repeats the paired comparison and tests whether the method effect changes across executions, separating directional reproducibility from effect-size stability. Applied across accuracy and execution metrics, REPAIR exposes three behaviors-directional reversal, magnitude shift, and effect attenuationand shows that reductions in reasoning turns and visible tool calls do not imply reproducible reductions in primitive computation or latency. The execution-level movement is comparable to, and here larger than, median gain reported by recent agentic VideoQA systems, contextualizing its magnitude without implying those systems are unstable. Significance within a single agentic execution is insufficient evidence that a reported method effect is reproducible.

视频问答可复现性评估协议

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。