arXiv:2511.02182cs.CV2025-11

提出触发帧机制,提升视频问答中的时空定位精度。

Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models

  • 用CORTEX提示生成关键帧作为视觉锚点。
  • 在GVQA任务上取得0.4968的HOTA分数,超越去年冠军。
  • 适合需要精准时空定位的多模态模型研究者。

本文针对ICCV 2025感知挑战赛中的视觉定位视频问答(GVQA)任务,提出一种三阶段框架:视频推理与问答、时空定位与跟踪。核心创新是通过自研CORTEX提示生成触发帧,即目标对象最显著的单帧,作为可靠的视觉锚点以支撑定位与跟踪。该方法在GVQA任务中实现0.4968的HOTA得分,显著优于去年冠军的0.2704,展现了在复杂视频内容推理与时空对齐方面的优越性能。

原文摘要 · Abstract (English)

In this technical report, we introduce a framework to address Grounded Video Question Answering (GVQA) task for the ICCV 2025 Perception Test Challenge. The GVQA task demands robust multimodal models capable of complex reasoning over video content, grounding the resulting answers visually, and tracking the referenced objects temporally. To achieve this capability, our proposed approach decomposes the GVQA task into a three-stage pipeline: (1) Video Reasoning \& QA, (2) Spatio-temporal Grounding and (3) Tracking. Our key contribution is the introduction of a trigger moment, derived from our proposed CORTEX prompt, which pinpoints the single most visible frame of a target object to serve as a robust anchor for grounding and tracking. To this end, we achieve the HOTA score of 0.4968, which marks a significant improvement over the previous year's winning score of 0.2704 on GVQA task.

视频问答时空定位多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。