用标记令牌统一图像与视频的区域理解,实现精准定位与跨模态关联。
Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
- 引入标记令牌,在视觉与文本中同步定位目标区域。
- 在RegVID-300k数据集上达到当前最优的区域级推理性能。
- 无需轨迹信息也能稳定理解视频区域,适合多模态任务研究者。
我们提出Omni-RGPT,一种面向图像与视频区域级理解的多模态大语言模型。为实现时空维度上一致的区域表示,引入Token Mark——一组用于标注视觉特征空间中目标区域的标记令牌。这些令牌通过区域提示(如框或掩码)直接嵌入空间区域,并同时融入文本提示以指定目标,建立视觉与文本令牌间的直接关联。为增强无需轨迹信息的视频理解鲁棒性,设计辅助任务,利用令牌一致性引导其稳定表达。此外,构建大规模区域级视频指令数据集RegVID-300k。Omni-RGPT在图像与视频常识推理基准上取得领先表现,同时在图像描述与指代表达理解任务中表现优异。
原文摘要 · Abstract (English)
We present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we introduce Token Mark, a set of tokens highlighting the target regions within the visual feature space. These tokens are directly embedded into spatial regions using region prompts (e.g., boxes or masks) and simultaneously incorporated into the text prompt to specify the target, establishing a direct connection between visual and text tokens. To further support robust video understanding without requiring tracklets, we introduce an auxiliary task that guides Token Mark by leveraging the consistency of the tokens, enabling stable region interpretation across the video. Additionally, we introduce a large-scale region-level video instruction dataset (RegVID-300k). Omni-RGPT achieves state-of-the-art results on image and video-based commonsense reasoning benchmarks while showing strong performance in captioning and referring expression comprehension tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。