让冻结的视频大模型主动找线索推理,无需训练或数据
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
- 分解问题+主动寻证+持续积累线索,三模块协同推理
- 在7个基准上平均提升5.9%,Egoschema上达10.5%
- 适合需要长时序推理但无法微调模型的场景
尽管大型视频语言模型(Video-LLMs)在感知任务上取得显著进展,其推理能力仍是瓶颈。现有方法通常依赖繁琐的数据工程——构建大规模思维链(CoT)数据集,再通过监督微调(SFT)和强化学习(RL)优化。该流程主要提升概率采样效率与输出分布对齐,却无法激发动态视觉探索所需的内在智能。本文提出TIR-Flow,一种新框架,将范式从被动处理转向主动视频搜索与推理,无需额外数据或参数更新。具体而言,框架包含三个协同模块:HDD将复杂查询分解为可验证子任务;HAP主动引导视觉注意力以获取高分辨率证据支持假设;EBA维持持久工作空间,持续累积并更新发现的线索以进行逻辑推理。在七个基准上的大量实验表明,TIR-Flow显著优于近期强基线,平均性能提升5.9%,在Egoschema上最高达10.5%。分析证实,赋予冻结的VLM类系统2级主动感知能力,是解决长时序视频推理的可扩展路径。
原文摘要 · Abstract (English)
While Large Video-Language Models (Video-LLMs) have achieved remarkable progress in perception, their reasoning capabilities remain a bottleneck. Existing solutions typically resort to a heavy "data engineering" paradigm-synthesizing large-scale Chain-of-Thought (CoT) datasets followed by Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). This pipeline primarily optimizes probability sampling efficiency and aligns output distributions, but fails to activate the intrinsic intelligence required for dynamic visual exploration. In this work, we propose TIR-Flow, a novel framework that shifts the paradigm from passive processing to active video searching and reasoning without additional data or parameter updating. Concretely, our framework operates through three synergistic modules: HDD decomposes complex queries into a set of verifiable sub-tasks; HAP actively directs visual attention to gather high-resolution evidence for hypothesis validation; EBA maintains a persistent workspace to accumulate and update the discovered clues for logical reasoning. Extensive experiments on seven benchmarks demonstrate that TIR-Flow significantly outperforms recent strong baselines, delivering an average performance boost of 5.9%, with gains reaching 10.5% on Egoschema. Our analysis confirms that empowering frozen VLMs with System-2-like active perception is a scalable path toward solving long-horizon video reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。