arXiv:2606.01104cs.CV2026-06

通过动态调整计算资源,提升视频问答中复杂关系推理的准确率。

Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge

论文配图:Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge
图 1 · 摘自论文原文
  • 先用轻量模型快速答题,再筛选不确定问题
  • 仅对难题启用高精度密集证据模块,实现90.07%平均准确率
  • 适合关注视频时序与空间关系推理的研究者

VRR-QA评估视频-语言系统是否能推断出单帧无法明确的空间、时间、视角、深度和可见性关系。我们提出一个仅依赖推理的系统,基于测试时自适应计算。系统首先通过视频-语言模型直接回答问题,然后利用多个轻量视图识别不稳定的题目。仅将这些困难问题传递给高预算的密集证据模块,该模块构建带时间戳的帧观测、特定关系探测器、候选验证及保守的时间聚合。此设计区分了两个常被混淆的问题:寻找可能的替代答案,以及判断当前答案是否应被更改。在测试集上,最终系统获得90.07%平均准确率和87.81%宏平均准确率。报告聚焦于最终测试系统的实现细节与复现所需设置。

原文摘要 · Abstract (English)

VRR-QA evaluates whether video-language systems can infer spatial, temporal, viewpoint, depth, and visibility relations that are not always resolved by a single frame. We present an inference-only system built around adaptive test-time computation. The system first answers each question with a direct video-language model pass, then uses multiple lightweight views to find unstable questions. Only these difficult questions are routed to a high-budget dense evidence module that constructs timestamped frame observations, relation-specific probes, candidate verification, and conservative temporal aggregation. This design separates two problems that are often confused in video question answering: finding plausible alternative answers and deciding when a current answer should actually be changed. On the test split, the final system obtains 90.07 average accuracy and 87.81 macro average accuracy. The report focuses on the final test system and the implementation settings required to reproduce the adaptive dense verifier.

视频问答关系推理自适应计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。