首个支持多粒度协同理解的视频大模型,能同时处理全局、像素和时间维度任务。
UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- 统一视觉语言对齐机制,跨尺度动态编码输入并输出文本、定位或掩码。
- 在三个协作任务上优于GPT-4o,9个公开基准验证了模型泛化能力。
- 适合需要多粒度视频分析的研究者与开发者使用。
随着多模态大语言模型的发展,视频大模型已在整体与专项视频理解任务中取得进展。然而,现有方法仅限于特定任务,难以实现全面且多层次的视频感知。为此,我们提出UFVideo,首个具备统一多粒度协同理解能力的视频大模型。通过设计统一的视觉-语言引导对齐机制,可在单一模型中灵活处理全局、像素与时间尺度的视频理解任务。UFVideo动态编码不同任务的视觉与文本输入,生成文本响应、时间定位或区域掩码。为评估挑战性的多粒度视频理解任务,我们构建了包含三个协作任务的UFVideo-Bench,结果表明其在灵活性与性能上优于GPT-4o。此外,我们在9个公开基准上验证了模型有效性,为未来视频大模型研究提供了重要参考。
原文摘要 · Abstract (English)
With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding tasks, failing to achieve a comprehensive and multi-grained video perception. To bridge this gap, we introduce UFVideo, the first Video LLM with unified multi-grained cooperative understanding capabilities. Specifically, we design unified visual-language guided alignment to flexibly handle video understanding across global, pixel and temporal scales within a single model. UFVideo dynamically encodes the visual and text inputs of different tasks and generates the textual response, temporal localization, or grounded mask. Additionally, to evaluate challenging multi-grained video understanding tasks, we construct the UFVideo-Bench consisting of three distinct collaborative tasks within the scales, which demonstrates UFVideo's flexibility and advantages over GPT-4o. Furthermore, we validate the effectiveness of our model across 9 public benchmarks covering various common video understanding tasks, providing valuable insights for future Video LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。