arXiv:2508.07330cs.CV2025-08中稿 · publication at ECA…

通过动态时空精修,让视频理解更懂复杂语言指令。

Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos

  • 用分句规划+逐步精修,让语言一步步引导视频特征优化
  • 在长复杂指令下准确率提升12.3%,优于当前最佳方法
  • 适合处理长查询、多动作链的视频定位与分割任务

视频-语言对齐需应对语言复杂性、动态交互实体、动作链及语义鸿沟。本文提出Planner-Refiner框架,通过迭代精修视觉元素的时空表征,以语言为指导缩小语义差距。规划模块将复杂语言提示分解为短句链;精修模块逐句处理名词-动词短语对,引导视觉标记在空间和时间上自注意力,实现单步高效精修。循环系统串联各步骤,保持优化后的视觉表示,并输入任务特定头生成对齐结果。在参考视频目标分割与时间定位两个任务上验证有效性,引入新基准MeViS-X评估长查询能力。在多个基准上性能超越现有最优方法,尤其在复杂提示下表现突出。

原文摘要 · Abstract (English)

Vision-language alignment in video must address the complexity of language, evolving interacting entities, their action chains, and semantic gaps between language and vision. This work introduces Planner-Refiner, a framework to overcome these challenges. Planner-Refiner bridges the semantic gap by iteratively refining visual elements' space-time representation, guided by language until semantic gaps are minimal. A Planner module schedules language guidance by decomposing complex linguistic prompts into short sentence chains. The Refiner processes each short sentence, a noun-phrase and verb-phrase pair, to direct visual tokens' self-attention across space then time, achieving efficient single-step refinement. A recurrent system chains these steps, maintaining refined visual token representations. The final representation feeds into task-specific heads for alignment generation. We demonstrate Planner-Refiner's effectiveness on two video-language alignment tasks: Referring Video Object Segmentation and Temporal Grounding with varying language complexity. We further introduce a new MeViS-X benchmark to assess models' capability with long queries. Superior performance versus state-of-the-art methods on these benchmarks shows the approach's potential, especially for complex prompts.

视频理解语言对齐时空建模长序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。