arXiv:2605.29402cs.CVcs.AI2026-05

分两步提取语义与视觉证据,提升长视频问答准确率

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

论文配图:Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
图 1 · 摘自论文原文
  • 用粗到细流程提取视频全局结构,用框和嵌入保留细节
  • 在多个任务上超越现有模型,关键在动态选择证据
  • 适合做长视频理解的开发者和研究者参考

长时视角视频理解对多模态大模型仍具挑战,受限于上下文长度和细粒度视觉信息的不足。新提出的HD-EPIC基准显示,即使强模型在多样化视频问答任务中表现也较弱。本文提出统一框架,将长视频推理解耦为语义证据与视觉证据两种互补形式:前者通过粗到细提取流程捕捉全局流程结构;后者以目标为中心,利用边界框和视觉嵌入保留细粒度定位。推理时,基于查询动态检索并融合两类证据。该方法在HD-EPIC-VQA挑战中多个任务类别取得竞争力表现。结果表明,显式构建、检索与整合语义和视觉证据,对多模态大模型实现有效长视频理解至关重要。

原文摘要 · Abstract (English)

Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insufficient grounding of fine-grained visual details. The recently proposed HD-EPIC benchmark highlights these limitations: even strong long-context models achieve relatively low performance across diverse video question answering tasks. In this paper, we propose a unified framework that decouples long-video reasoning into two complementary forms of evidence: semantic evidence and visual evidence. Semantic evidence captures global procedural structure through a coarse-to-fine extraction pipeline, while object-centric visual evidence preserves fine-grained grounding through bounding boxes and visual embeddings. During inference, we formulate reasoning as a query-conditioned evidence retrieval and integration process, dynamically selecting relevant information from both sources. Our approach achieves competitive performance in the HD-EPIC-VQA Challenge across multiple task categories. More broadly, our results demonstrate that explicitly structuring, retrieving, and integrating semantic and visual evidence is critical for effective long-video understanding with MLLMs.

视频理解多模态长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。