构建多轮视频对话系统,实现精准定位与语义理解的统一。
SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
- 设计联合学习框架,整合视频指代理解与定位能力。
- 在522个视频上构建5067个问题的评估基准,性能领先。
- 适合研究多轮视频交互、视觉语言模型的学者使用。
当前视频大模型在细粒度时空理解方面仍面临挑战,关键在于缺乏高质量的统一指令数据和全面评估基准。本文提出SAMA,包含三大贡献:一是构建SAMA-239K数据集,涵盖15,000个视频,支持视频指代理解、定位与多轮对话联合训练;二是提出SAMA模型,融合时空上下文聚合器与Segment Anything Model,提升细粒度视频理解与精准定位能力;三是建立SAMA-Bench基准,包含5,067个来自522个视频的问题,全面评估视频大模型在多轮时空指代与基于视觉的对话中的综合表现。实验表明,SAMA不仅在SAMA-Bench上表现优异,还在通用定位基准上达到新SOTA,并在标准视觉理解任务中保持竞争力。
原文摘要 · Abstract (English)
Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring understanding, which captures the semantics of video regions, and video grounding, which segments object regions based on natural language descriptions. However, most existing approaches tackle these tasks in isolation, limiting progress toward unified, referentially grounded video interaction. We identify a key bottleneck in the lack of high-quality, unified video instruction data and a comprehensive benchmark for evaluating referentially grounded video chat. To address these challenges, we contribute in three core aspects: dataset, model, and benchmark. First, we introduce SAMA-239K, a large-scale dataset comprising 15K videos specifically curated to enable joint learning of video referring understanding, grounding, and multi-turn video chat. Second, we propose the SAMA model, which incorporates a versatile spatio-temporal context aggregator and a Segment Anything Model to jointly enhance fine-grained video comprehension and precise grounding capabilities. Finally, we establish SAMA-Bench, a meticulously designed benchmark consisting of 5,067 questions from 522 videos, to comprehensively evaluate the integrated capabilities of Video LMMs in multi-turn, spatio-temporal referring understanding and grounded dialogue. Extensive experiments and benchmarking results show that SAMA not only achieves strong performance on SAMA-Bench but also sets a new state-of-the-art on general grounding benchmarks, while maintaining highly competitive performance on standard visual understanding benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。