arXiv:2607.02551cs.CVcs.AI2026-07

用视频间差异训练模型,提升细粒度时空感知能力。

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences

论文配图:DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
图 1 · 摘自论文原文
  • 通过对比相似视频中的局部差异,让模型学习定位变化
  • 在10K样本数据集上提升跨视频差异理解性能,泛化至多个基准
  • 适合关注视频细节推理与模型可解释性的研究者

视频多模态大语言模型在开放域视频理解上取得显著进展,但仍缺乏精确的局部时空感知能力。当两段视频整体语义几乎相同,仅在短时间或小区域存在差异时,现有模型常无法发现变化并提供可靠证据。本文提出DELTAVID,一种基于跨视频差异的可验证代理任务框架,通过将“找不同”转化为可训练的感知信号,使模型能识别局部变化、判断时间边界、组织空间证据。为实现可扩展训练与可靠评估,构建了DELTAVID-10K和DELTAVID-Bench,将真实视频中可控的局部差异转化为带证据标注的训练与测试样本。实验表明,DELTAVID显著提升跨视频差异理解性能,并将学到的局部证据能力迁移至多个通用视频理解基准,包括MMVU、MLVU、Video-MME、VideoHolmes、VideoMMMU、LVBench、TempCompass和LongVideoBench。结果证明,跨视频差异不仅是诊断细粒度感知失败的有效方式,更是一种可扩展的代理监督信号,推动视频多模态大模型从粗粒度语义理解迈向细粒度时空证据推理。

原文摘要 · Abstract (English)

Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception. When two videos share almost the same global semantics and differ only in a short time span or a small region, current models often fail to find the change and provide reliable evidence. We propose DELTAVID, a verifiable proxy-task framework that enhances fine-grained spatiotemporal perception with cross-video differences. The key idea is to turn cross-video spot-the-difference into a trainable perception signal, where a model identifies local changes, judges temporal boundaries, and organizes spatial evidence by comparing similar videos. To make this signal scalable to train and reliable to evaluate, we further introduce DELTAVID-10K and DELTAVID-Bench, which convert controllable local differences in real videos into evidence-labeled training and test samples. Experiments show that DELTAVID substantially improves performance on cross-video difference understanding and transfers the learned local evidence ability to general video understanding benchmarks, including MMVU, MLVU, Video-MME, VideoHolmes, VideoMMMU, LVBench, TempCompass, and LongVideoBench. These results show that cross-video differences are not only an effective way to diagnose fine-grained perception failures, but also a scalable proxy supervision that moves Video MLLMs from coarse semantic understanding toward fine-grained spatiotemporal evidence reasoning.

视频理解细粒度感知多模态证据推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。