提出可验证证据充分性的智能体,提升长视频问答的准确性
REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering

- 用视觉相似性自适应分组帧,构建动态视频记忆
- 通过评分标准库显式检查证据是否充分,缺失则重查
- 无需训练即可超越现有方法,适合长视频推理场景
近期,检索增强和记忆增强方法成为长视频问答的两大主流范式。然而,现有方法通常依赖固定时长(如10秒)的时间切片和静态离线记忆库,不仅割裂连贯事件,且无法实时适应推理需求。此外,无论采用多尺度摘要还是多模态知识图谱,当前方法均侧重检索相关性而忽视证据充分性,一旦获得语义相关的线索便停止回答,即使关键的时间、因果或细粒度动作证据仍缺失。为此,我们提出REVEAL——一种基于评分标准的智能体框架。其基础为自适应视觉相似性预处理流程,将视觉上连续的相邻帧聚合成自然事件单元,构建离线-在线结合的视频记忆:离线捕捉全局上下文,线上动态维护与问题相关的记忆。在此结构化记忆基础上,REVEAL利用自动生成的评分标准库,显式验证检索证据是否满足充分性要求,失败时定位缺失线索并引导针对性重检索。无需额外训练,REVEAL在多项实验中持续优于闭源与开源的最先进方法。结果表明,显式验证证据充分性而非仅依赖语义相关性,能找回先前方法遗漏的关键线索,实现更可靠的长视频推理。
原文摘要 · Abstract (English)
Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. However, existing methods typically rely on rigid, fixed-length temporal chunking (e.g., 10s) and static offline memory banks, which not only fragment coherent continuous events but also fail to adapt during real-time reasoning. Moreover, whether using multi-scale summaries or multimodal knowledge graphs, current approaches prioritize retrieval relevance while overlooking evidence sufficiency, often stopping to answer once only semantically relevant clues are retrieved, even when key temporal, causal, or fine-grained action evidence is still missing. To tackle these challenges, we propose REVEAL, a rubric-guided agent framework. As a foundation, we introduce an adaptive visual-similarity-based preprocessing pipeline that groups visually coherent adjacent frames into natural event units to construct an offline-online video memory---capturing global video context offline while dynamically maintaining question-conditioned memory online. Built upon this structured memory, REVEAL uses an automatically constructed rubric library to explicitly verify whether retrieved evidence satisfies sufficiency criteria, pinpoints missing clues upon verification failure, and directs targeted re-retrieval for complementary information. Without any extra training, REVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments. These results show that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。