用视频间差异训练模型,提升细粒度时空感知能力。
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

- 通过对比相似视频中的局部差异,让模型学习定位变化
- 在10K样本数据集上提升跨视频差异理解性能,泛化至多个基准
- 适合关注视频细节推理与模型可解释性的研究者
视频多模态大语言模型在开放域视频理解上取得显著进展,但仍缺乏精确的局部时空感知能力。当两段视频整体语义几乎相同,仅在短时间或小区域存在差异时,现有模型常无法发现变化并提供可靠证据。本文提出DELTAVID,一种基于跨视频差异的可验证代理任务框架,通过将“找不同”转化为可训练的感知信号,使模型能识别局部变化、判断时间边界、组织空间证据。为实现可扩展训练与可靠评估,构建了DELTAVID-10K和DELTAVID-Bench,将真实视频中可控的局部差异转化为带证据标注的训练与测试样本。实验表明,DELTAVID显著提升跨视频差异理解性能,并将学到的局部证据能力迁移至多个通用视频理解基准,包括MMVU、MLVU、Video-MME、VideoHolmes、VideoMMMU、LVBench、TempCompass和LongVideoBench。结果证明,跨视频差异不仅是诊断细粒度感知失败的有效方式,更是一种可扩展的代理监督信号,推动视频多模态大模型从粗粒度语义理解迈向细粒度时空证据推理。
原文摘要 · Abstract (English)
Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and differ only in a short time span or a small region, current models often fail to find the change and provide reliable evidence. We propose DELTAVID, a verifiable proxy-task framework that enhances fine-grained spatiotemporal perception with cross-video differences. The key idea is to turn cross-video spot-the-difference into a trainable perception signal, where a model identifies local changes, judges temporal boundaries, and organizes spatial evidence by comparing similar videos. To make this signal scalable to train and reliable to evaluate, we further introduce DELTAVID-10K and DELTAVID-Bench, which convert controllable local differences in real videos into evidence-labeled training and test samples. Experiments show that DELTAVID substantially improves performance on cross-video difference understanding and transfers the learned local evidence ability to general video understanding benchmarks, including MMVU, MLVU, Video-MME, VideoHolmes, VideoMMMU, LVBench, TempCompass, and LongVideoBench. These results show that cross-video differences are not only an effective way to diagnose fine-grained perception failures, but also a scalable proxy supervision that moves Video MLLMs from coarse semantic understanding toward fine-grained spatiotemporal evidence reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。