arXiv:2608.10764cs.CV2026-08

让AI从猜答案变成主动找证据,提升视频理解的推理能力。

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

论文配图:FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding
图 1 · 摘自论文原文
  • 先让模型聚焦视觉异常,再逐步去掉文字提示,逼它自主发现证据。
  • 在三个任务上超越GPT-5.6,开放问答和描述任务性能保留率分别达90.4%和67.4%。
  • 可将现有题目自动转为新任务,无需额外标注,适合研究视频因果推理。

反事实视频理解旨在检验模型是否掌握物理与常识规律。然而,现有基于多选题(MCQ)的评测基准会通过问题和选项泄露目标事件,使挑战从主动发现变为文本引导验证。本文提出FADE框架,实现反事实发现与解释的有效训练。方法采用证据优先、两阶段范式:首先通过证据内化监督微调,使模型预测基于决定性视觉异常;其次应用渐退锚点强化学习策略,逐步移除文本引导,迫使模型独立发现并解释证据。为严格评估该能力,我们还设计了有效管道,将现有MCQ数据集转换为对齐的MCQ、开放问答(OQA)和字幕生成任务,无需额外数据标注。以Qwen3-VL-8B为基线,FADE在DualityVidQA-test和IPV-Bench上三项任务均达最佳严格配对得分,优于GPT-5.6。尤其在从约束性多选题转向开放问答和字幕生成时,模型表现保持力显著更强:性能保留率分别为90.4%和67.4%,远超GPT-5.6的48.1%和30.7%。我们希望此简单框架能成为未来无约束反事实视频理解研究的坚实基线。

原文摘要 · Abstract (English)

Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ) benchmarks inadvertently leak target events through their questions and candidate options. This reduces the core challenge from active discovery to text-guided verification. In this paper, we present FADE, an effective training framework for counterfactual discovery and explanation. Our method is built on an evidence-first, two-stage training paradigm. First, evidence-internalized supervised fine-tuning grounds the model's predictions in decisive visual anomalies. Second, we apply a fading-anchor reinforcement learning strategy that progressively removes textual guidance, compelling the model to independently discover and explain evidence. To rigorously evaluate this capability, we also introduce an effective pipeline that converts existing MCQ datasets into aligned MCQ, open-ended question answering (OQA), and captioning tasks without requiring additional data curation. Our simple approach yields strong results. Using Qwen3-VL-8B as the baseline, FADE achieves state-of-the-art strict paired scores across all three tasks on DualityVidQA-test and IPV-Bench, outperforming GPT-5.6. In specific, when transitioning from constrained MCQs to unconstrained OQA and captioning, our model demonstrates remarkable robustness. Its performance retention is 90.4% and 67.4% on DualityVidQA-test-substantially higher than the 48.1% and 30.7% retained by GPT-5.6. We hope this simple framework can serve as a solid baseline for future research in unconstrained counterfactual video understanding.

视频理解反事实推理自监督学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。