让视频事件定位更准,靠的是整体感知而非单帧匹配。
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
- 用特殊标记整合事件所有帧信息,保持语义连贯性。
- 通过平滑处理提升时间匹配精度,减少噪声干扰。
- 多粒度特征聚合弥补压缩损失,适合细粒度视频理解任务。
尽管视频大语言模型(Vid-LLMs)取得进展,但时间视频定位(TVG)仍具挑战性。现有方法通常通过比较起止帧特征与两个独立标记进行匹配,依赖精确时间戳,难以捕捉事件的语义连续性和完整性,导致歧义。为此,我们提出E.M.Ground,一种专注于整体事件感知的新型Vid-LLM。其核心创新包括:(i) 引入<event>特殊标记,聚合查询事件所有帧的信息,保留语义连续性以实现精准匹配;(ii) 采用Savitzky-Golay平滑处理帧间相似度噪声,提升预测准确性;(iii) 多粒度帧特征聚合,增强匹配可靠性并缓解压缩带来的信息损失。在基准数据集上的大量实验表明,E.M.Ground显著优于现有最优Vid-LLMs。
原文摘要 · Abstract (English)
Despite recent advances in Video Large Language Models (Vid-LLMs), Temporal Video Grounding (TVG), which aims to precisely localize time segments corresponding to query events, remains a significant challenge. Existing methods often match start and end frames by comparing frame features with two separate tokens, relying heavily on exact timestamps. However, this approach fails to capture the event's semantic continuity and integrity, leading to ambiguities. To address this, we propose E.M.Ground, a novel Vid-LLM for TVG that focuses on holistic and coherent event perception. E.M.Ground introduces three key innovations: (i) a special <event> token that aggregates information from all frames of a query event, preserving semantic continuity for accurate event matching; (ii) Savitzky-Golay smoothing to reduce noise in token-to-frame similarities across timestamps, improving prediction accuracy; (iii) multi-grained frame feature aggregation to enhance matching reliability and temporal understanding, compensating for compression-induced information loss. Extensive experiments on benchmark datasets show that E.M.Ground consistently outperforms state-of-the-art Vid-LLMs by significant margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。