arXiv:2606.25585cs.CV2026-06中稿 · ECCV

让模型预测未来动作并回溯当前对象,突破传统视频分割局限。

FeVOS: Foresight Expression Video Object Segmentation

论文配图:FeVOS: Foresight Expression Video Object Segmentation
图 1 · 摘自论文原文
  • 设计新任务:基于未来事件提问,要求回溯当前对象掩码。
  • 构建968段视频数据集,含1.45万条前瞻表达与2904条推理链。
  • 提出多模态大模型FeVOS-R1,可跨任务泛化,适合未来感知研究者。

现有指代视频对象分割任务主要关注对已观测帧中物体事件、动作或外观的描述,缺乏对需预先决策的时空推理场景的评估,限制了其应用范围。为此,我们提出前瞻表达视频对象分割任务,要求模型根据未来视频片段中的事件提问,并以当前帧中对象的掩码作为视觉回答。例如,在第一人称视角场景中,问题“将使用什么工具?”需结合时空线索预测下一将使用的工具掩码,有助于理解未来动作与决策。为此,我们构建了包含968个视频片段、14,525条前瞻表达和2,904条思维链标注的FeVOS数据集,提供明确且可解释的推理步骤。我们进一步开发了基于多模态大模型的FeVOS-R1模型,通过监督微调与强化学习两阶段训练。FeVOS-R1不仅在FeVOS上达到当前最优性能,还展现出对现有RVOS基准的强大泛化能力。我们希望本工作能激发更多关于视频感知中预测性推理的研究。

原文摘要 · Abstract (English)

Existing Referring Video Object Segmentation tasks focus on referring expressions describing events, actions or appearances of relevant objects within the observed frames, lacking evaluation in scenarios that require pre-decisive spatio-temporal reasoning, thereby limiting their applicability. To address this, we propose Foresight Expression Video Object Segmentation, a task that queries future events in upcoming video segments and requires masks of the objects in the observed frames as visual answers. For example, in ego-centric scenes, the question "What tool will be used?" demands reasoning over spatio-temporal cues to predict the masks of the next tool to be used, which helps with the understanding of future actions and decisions. To support this task, we introduce FeVOS, a dataset with 968 video clips, 14,525 foresight expressions, and 2,904 chain-of-thought annotations to provide explicit and interpretable reasoning steps. We further develop FeVOS-R1, an MLLM-based model trained on our dataset via a two-stage pipeline of supervised fine-tuning and reinforcement learning. FeVOS-R1 not only achieves state-of-the-art performance on FeVOS, but also demonstrates strong generalization to existing RVOS benchmarks. We hope this work can inspire more research on predictive reasoning in video perception.

视频分割前瞻推理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。