数据集结构决定视频模型设计,揭示了模型演进的深层逻辑。
Video Understanding by Design: How Datasets Shape Video Models
- 从数据集结构出发,解析模型架构的演化动因
- 不同数据需求催生特定归纳偏置,如时序敏感或跨模态对齐
- 适合关注模型设计原理与数据驱动创新的研究者
视频理解研究快速发展,得益于日益多样的数据集和更强大的模型架构。现有综述多按任务、基准或模型家族组织进展,却难以解释为何特定架构出现并成功。本文提出以数据集为中心的视角,将数据集结构、归纳偏置与架构设计统一于一个框架中。我们指出,不同数据集要求模型具备特定不变性与能力,如对视角变化的鲁棒性、对时序顺序的敏感性、长程依赖推理、关系交互建模及跨模态对齐。这些需求自然产生归纳偏置——即偏好特定推理与泛化模式的架构假设。在此视角下,两流网络、3D CNN、时序模型、Transformer、图方法及多模态基础模型等里程碑架构,均可视为对不断演化的数据集挑战的回应。基于此框架,我们系统分析了数据集特征如何塑造各任务中的架构创新,并讨论了不同数据范式带来的表征偏置。该综述不仅回顾了领域演进,也为通用视频理解系统提供了前瞻性路线图。代码与动态视频可视化见 https://time.griffith.edu.au/paper-sites/video-understanding/。
原文摘要 · Abstract (English)
Research in video understanding has advanced rapidly, driven by increasingly diverse datasets and more powerful model architectures. While existing surveys typically organize progress by tasks, benchmarks, or model families, they provide limited insight into why particular architectures emerged and succeeded. In this survey, we argue that the evolution of video understanding is fundamentally shaped by dataset structure. We present a dataset-centric perspective that connects dataset structure, inductive biases, and architectural design within a unified framework. We show that different datasets require models to capture specific invariances and capabilities, such as robustness to viewpoint changes, sensitivity to temporal ordering, reasoning over long-range dependencies, relational interactions, and cross-modal alignment. These requirements naturally give rise to inductive biases, i.e., architectural assumptions that favor particular patterns of reasoning and generalization. From this perspective, milestone architectures, including two-stream networks, 3D CNNs, temporal models, transformers, graph-based methods, and multimodal foundation models, can be understood as architectural responses to the challenges posed by evolving datasets. Building on this framework, we systematically analyze how dataset characteristics have shaped architectural innovation across video understanding tasks and discuss the representational biases induced by different data regimes. By unifying datasets, inductive biases, and architectures into a coherent perspective, this survey offers both a retrospective explanation of the field's evolution and a forward-looking roadmap toward general-purpose video understanding systems. Code and dynamic video visualizations of dataset-induced biases are available at https://time.griffith.edu.au/paper-sites/video-understanding/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。