arXiv:2608.16332cs.CV2026-08中稿 · ACM MM2026

让语言描述中的动作语义显式驱动视频目标分割,提升定位精度

Unlocking Motion in Expressions: Temporal Calibration for Referring Video Object Segmentation

论文配图:Unlocking Motion in Expressions: Temporal Calibration for Referring Video Object Segmentation
图 1 · 摘自论文原文
  • 从语言描述中提取可解释的动作控制信号,显式建模动作语义
  • 在6个标准数据集上均超越现有方法,最高提升达3.2% mIoU
  • 适合需要精准时序理解的视频理解与交互应用

指代视频目标分割(RVOS)旨在根据自然语言描述,在视频序列中进行像素级目标分割。现有方法通常将运动信息整合进统一的跨模态时序建模框架中,用语言线索实现目标定位与分割。然而,表达对运动语义的依赖未被显式建模,导致难以根据不同语义需求自适应调整运动信息的使用。为此,本文提出表达驱动的运动校准(EMC)框架,显式挖掘并利用语言描述中的运动语义。该方法通过运动信号处理(MSP)模块从表达中提取可解释的运动控制信号,并通过运动影响校准(MIC)模块在时序决策中动态调节运动线索的贡献。此外,引入语义时序阶段构建(STSC)模块,建立与表达相关的时序阶段,为运动校准提供紧凑的候选时序空间。在六个标准基准测试(包括Ref-YouTubeVOS、Ref-DAVIS17、MeViS(valid/valid$^u$)、A2D-Sentences和JHMDB-Sentences)上进行了广泛评估,验证了方法的优越性。代码将发布于https://github.com/Jeven7/EMC。

原文摘要 · Abstract (English)

Referring Video Object Segmentation (RVOS) aims to segment referred objects at the pixel level in video sequences based on natural language descriptions. Existing methods typically introduce motion information within a unified cross-modal temporal modeling framework, where language cues are used for target localization and segmentation. However, the dependency of expressions on motion semantics is not explicitly modeled, making it difficult to adaptively adjust the use of motion information according to different semantic requirements. To address these issues, we propose an Expression-driven Motion Calibration (EMC) framework for RVOS that explicitly unlocks and leverages the motion semantics within expressions. The proposed method extracts interpretable motion control signals from expressions via a Motion Signal Processing (MSP) module, and employs a Motion Influence Calibration (MIC) module to adjust the contribution of motion cues during temporal decision making. In addition, a Semantic Temporal Stage Construction (STSC) module is introduced to build expression-relevant temporal stages, providing a compact temporal candidate space for motion calibration. Through extensive evaluation on six standard benchmarks, including Ref-YouTubeVOS, Ref-DAVIS17, MeViS (valid/valid$^u$), A2D-Sentences, and JHMDB-Sentences, the superiority of our method is validated. We will release the code on https://github.com/Jeven7/EMC.

视频分割时序建模多模态语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。