arXiv:2507.03578cs.CVcs.AI2025-07ICCV被引 5

构建跨科学领域的视频模型评测基准,验证通用模型在医学、动物行为等场景的迁移能力。

SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications

  • 设计五个跨领域的科学视频任务,评估通用视频模型的迁移性能
  • 仅用可训练读出模块适配,即在多个任务达到顶尖水平
  • 揭示现有模型局限,推动更通用科学视频模型发展

近年来,各类时空基础模型在不同科学领域迅速发展。尽管前景广阔,这些模型通常具有领域特异性,且仅在设计任务中进行评估。鉴于许多科学任务可建模为视频问题,视频基础模型(ViFMs)有望成为通用、无领域依赖的解决方案。然而,大规模但可能非目标领域的预训练知识能否有效跨学科迁移,以及单一预训练ViFM是否能媲美领域专用基线,仍不明确。为此,我们提出SciVid,一个涵盖医学计算机视觉、动物行为分析和天气预报的综合性基准,包含五个科学视频任务。通过简单可训练读出模块适配六种领先ViFMs,我们建立了强基线,并证明了通用表示的有效迁移潜力。具体而言,在多个应用中仅靠预训练骨架即可实现当前最优结果。此外,我们的结果揭示了现有ViFMs的局限性,指出了面向高影响力科学应用的通用化模型研发机遇。代码已开源:https://github.com/google-deepmind/scivid。

原文摘要 · Abstract (English)

In recent years, there has been a proliferation of spatiotemporal foundation models in different scientific disciplines. While promising, these models are often domain-specific and are only assessed within the particular applications for which they are designed. Given that many tasks can be represented as video modeling problems, video foundation models (ViFMs) hold considerable promise as general-purpose domain-agnostic approaches. However, it is not known whether the knowledge acquired on large-scale but potentially out-of-domain data can be effectively transferred across diverse scientific disciplines, and if a single, pretrained ViFM can be competitive with domain-specific baselines. To address this, we introduce SciVid, a comprehensive benchmark comprising five *Sci*entific *Vid*eo tasks, across medical computer vision, animal behavior, and weather forecasting. We adapt six leading ViFMs to SciVid using simple trainable readout modules, establishing strong baselines and demonstrating the potential for effective transfer learning. Specifically, we show that state-of-the-art results can be obtained in several applications by leveraging the general-purpose representations from ViFM backbones. Furthermore, our results reveal the limitations of existing ViFMs, and highlight opportunities for the development of generalizable models for high-impact scientific applications. We release our code at https://github.com/google-deepmind/scivid to facilitate further research in the development of ViFMs.

视频模型科学应用迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。