通过结构化提示提升视频模型时空推理能力
Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

- 在推理时添加轻量级时空结构,帮助模型更好组织视觉证据
- 在多个视频任务中实现性能提升,具体增益因模型和任务而异
- 无需训练、不改模型,适合快速提升现有视频理解系统
视频-语言模型(VLMs)在需要跨时间追踪事件并定位到特定空间区域的任务上仍表现脆弱。我们提出,部分限制可通过推理时更优的视觉证据组织方式缓解。本文引入结构化视频提示(structured video prompting),一种无需训练的推理时方法,在输入视频中加入轻量级的空间与时间结构,为跨时空组织证据提供显式锚点,不改变模型权重或解码过程,也不修改问题提示。我们在两个互补的视频基准测试和两个开放视频语言模型上评估该方法。结果表明,在多种设置下,结构化输入可带来性能提升,增益程度随模型和任务而异。研究发现,部分VLM失败不仅源于推理能力不足,也与推理时视频证据的呈现方式有关。这些结果强调了结构化视频提示作为改进视频理解的简单且实用方向。
原文摘要 · Abstract (English)
Video-language models (VLMs) remain brittle on tasks that require tracking events over time and grounding answers in specific spatial regions. We propose that part of this limitation can be addressed through better organization of visual evidence at inference time. We introduce structured video prompting, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding and without altering the question prompt in the main comparison. We evaluate this approach on two complementary video benchmarks and two open video-language models. Across these settings, structured inputs improve performance in several cases, with gains varying by model and task. Our findings suggest that some failures of VLMs arise not only from reasoning capacity, but also from how video evidence is presented at inference time. These results highlight structured video prompting as a simple and practical direction for improving video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。