用合成数据让视频大模型学会精准定位时空信息。
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
- 自动生成带时空标记的指令数据,模拟真实场景中的指代问题。
- 在无需人工标注的情况下,显著提升模型对时空指代的理解能力。
- 适合需要精确理解视频中物体位置与时间事件的应用场景。
下一代AI助手需超越通用视频理解,精准解析动态现实环境中的空间与时间指代。现有视频大语言模型虽具备粗粒度理解能力,但在依赖时间事件锚定或手势线索进行空间定位时表现不佳。为此,我们提出Strefer,一种用于生成合成指令数据的框架,通过伪标注高密度、细粒度的视频元数据,结构化地捕捉主体、对象、其位置(masklets)、动作描述及时间线等丰富信息。该方法增强了视频大模型对空间与时间指代的解析能力,推动更灵活、时空感知更强的推理发展。实验表明,在不使用私有模型、昂贵人工标注或大规模新视频标注的前提下,基于Strefer数据训练的模型在时空消歧任务上优于基线,并展现出更强的时空感知推理能力,为感知基础的指令微调视频大模型奠定了新基础。
原文摘要 · Abstract (English)
Next-generation AI companions must go beyond general video understanding to resolve spatial and temporal references in dynamic, real-world environments. Existing Video Large Language Models (Video LLMs), while capable of coarse-level comprehension, struggle with fine-grained, spatiotemporal reasoning, especially when user queries rely on time-based event references for temporal anchoring, or gestural cues for spatial anchoring to clarify object references and positions. To bridge this critical gap, we introduce Strefer, a synthetic instruction data generation framework designed to equip Video LLMs with spatiotemporal referring and reasoning capabilities. Strefer produces diverse instruction-tuning data using a data engine that pseudo-annotates temporally dense, fine-grained video metadata, capturing rich spatial and temporal information in a structured manner, including subjects, objects, their locations as masklets, and their action descriptions and timelines. Our approach enhances the ability of Video LLMs to interpret spatial and temporal references, fostering more versatile, space-time-aware reasoning essential for real-world AI companions. Without using proprietary models, costly human annotation, or the need to annotate large volumes of new videos, experimental evaluations show that models trained with data produced by Strefer outperform baselines on tasks requiring spatial and temporal disambiguation. Additionally, these models exhibit enhanced space-time-aware reasoning, establishing a new foundation for perceptually grounded, instruction-tuned Video LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。