用证据账本增强视频关系推理,提升问答准确率。
Question-Aware Evidence Ledgers for Video Relational Reasoning

- 构建问题感知的证据账本,动态提取关键信息
- 在测试集上达到92.95%整体准确率,93.79%宏平均准确率
- 适合需要精准时空与对话理解的任务场景
VRR-QA挑战评估视频中的视觉关系推理能力,答案常依赖隐含的空间关系、事件边界、目标身份和对话上下文,而非单一显著帧。本文提出一种基于强GPT-5.5视频问答模型的测试时推理流水线,结合一组问题感知的证据账本。初始模型从统一视频表示中回答问题,而路由账本被提示显式化目标、计数单位、参考帧、时间或空间范围,以支持计数、空间、端点、视角和对话推理。外部工具如开放词汇检测、深度线索、成对裁剪、语音识别(ASR)和场景图账本仅作为证据来源。保守门控机制确保当前答案保留,除非独立证据唯一支持另一选项。最终证据门控流水线在挑战测试集上取得92.95%的整体准确率和93.79%的宏平均准确率。
原文摘要 · Abstract (English)
The VRR-QA challenge evaluates visual relational reasoning in videos, where answers often depend on implicit spatial relations, event boundaries, target identity, and dialogue context rather than a single salient frame. We present a test-time reasoning pipeline built around a strong GPT-5.5 video QA solver and a set of question-aware evidence ledgers. The initial solver answers each question from a uniform video representation, while routed ledgers are prompted to make the required targets, count units, reference frames, and temporal or spatial scope explicit for counting, spatial, endpoint, viewpoint, and dialogue reasoning. External tools such as open-vocabulary detection, depth cues, pair crops, ASR, and scene-graph ledgers are used only as evidence sources. A conservative gate keeps the current answer unless independent evidence uniquely supports a different option. The final evidence-gated pipeline achieves 92.95% overall accuracy and 93.79% macro accuracy on the challenge test split.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。