提出ClueNet框架,用视觉线索提升视频问答的准确性与可解释性。
Clue Matters: Leveraging Latent Visual Clues to Empower Video Reasoning
- 分两阶段训练,分离线索提取与链式推理,提升结构化理解能力。
- 在NExT-QA等数据集上性能领先1.1%以上,显著减少幻觉现象。
- 适合高要求视频问答场景,支持多种模型架构且推理高效。
多模态大语言模型虽推动了视频理解发展,但视频问答仍面临时间因果推理和证据支撑答案生成的挑战。现有端到端框架缺乏视觉感知与答案推导间的显式结构化推理,导致严重幻觉和可解释性差。现有方法未能解决三大核心问题:忠实的视觉线索提取、基于效用的线索筛选、端到端线索-答案对齐。受人类视觉认知层级启发,我们提出ClueNet,一种线索感知的视频推理框架,采用两阶段监督微调,无需大规模基础模型修改。解耦监督实现线索提取与链式推理对齐,推理阶段通过自适应线索过滤器优化高阶推理,辅以轻量模块提升推理效率。在NExT-QA、STAR和MVBench上的实验表明,ClueNet性能优于当前最优方法至少1.1%,具备更强泛化能力、更少幻觉、更高推理效率及跨骨干网络兼容性。该工作弥合了多模态大模型视频理解中从感知到生成的鸿沟,为高风险视频问答应用提供可解释、忠实的推理范式。
原文摘要 · Abstract (English)
Multi-modal Large Language Models (MLLMs) have significantly advanced video reasoning, yet Video Question Answering (VideoQA) remains challenging due to its demand for temporal causal reasoning and evidence-grounded answer generation. Prevailing end-to-end MLLM frameworks lack explicit structured reasoning between visual perception and answer derivation, causing severe hallucinations and poor interpretability. Existing methods also fail to address three core gaps: faithful visual clue extraction, utility-aware clue filtering, and end-to-end clue-answer alignment. Inspired by hierarchical human visual cognition, we propose ClueNet, a clue-aware video reasoning framework with a two-stage supervised fine-tuning paradigm without extensive base model modifications. Decoupled supervision aligns clue extraction and chain-based reasoning, while inference supervision with an adaptive clue filter refines high-order reasoning, alongside lightweight modules for efficient inference. Experiments on NExT-QA, STAR, and MVBench show that ClueNet outperforms state-of-the-art methods by $\ge$ 1.1%, with superior generalization, hallucination mitigation, inference efficiency, and cross-backbone compatibility. This work bridges the perception-to-generation gap in MLLM video understanding, providing an interpretable, faithful reasoning paradigm for high-stakes VideoQA applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。