arXiv:2606.26994cs.CVcs.AI2026-06被引 1

将视频拆解为事件单元,提升指代视频分割的准确性

Event-Aware Instructed Assistant for Referring Video Segmentation

论文配图:Event-Aware Instructed Assistant for Referring Video Segmentation
图 1 · 摘自论文原文
  • 用可学习的事件查询分解视频为多个简单事件
  • 在5个公开数据集上实现领先性能,显著减少混淆和幻觉
  • 适合需要精准理解复杂视频语义的任务场景

现有指代视频分割方法通常将视频视为由多帧图像组成的单一事件,忽视了视频中包含多个独立事件的事实。这种机制要求模型直接理解复杂的视频与文本内容,易导致混淆和幻觉。为此,我们提出通过可学习的事件查询将视频分解为一组简单事件,以逐事件、易理解的方式解析复杂内容。基于自然语言表达常将视频划分为与文本相关的独立片段这一观察,我们引入EVIS(Event-Aware Video Instructed Segmentation Assistant),利用文本引导的事件查询将视频划分为简单事件,并提取事件感知的视觉-文本特征,实现视频的分层理解。此外,我们提出对象-像素混合学习(Object-Pixel-Hybrid Learning),通过融合细粒度像素特征与先验对象查询,使多模态大模型在长视频中实现目标追踪。在5个公开基准上的大量实验表明,EVIS在指代视频分割任务上表现优异。

原文摘要 · Abstract (English)

Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact that a video typically contains multiple distinct events. Under such a mechanism, the model needs to directly understand all the complex content in the video and text, which can easily lead to confusion and hallucinations. To address this issue, we propose to decompose a video to a set of simple events by learnable Event Query, and understand complex video content in an event-by-event, easy-to-understand manner. This is based on the observation that natural language expressions often divide a video into distinct, text-related segments, each representing a separate event within a compound event. We introduce EVIS, an Event-Aware Video Instructed Segmentation Assistant, which utilizes text-guided Event Queries to partition a video into simple events, extracting event-aware visual-text features to achieve a hierarchical understanding of the video. Additionally, we propose Object-Pixel-Hybrid Learning, which enables the MLLMs to track targets in long-term videos by integrating fine-grained pixel features with prior object queries. Extensive experimental results on 5 public benchmarks demonstrate EVIS's strong performance in addressing the referring video segmentation task.

视频分割事件理解多模态长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。