arXiv:2601.06287cs.CV2026-01

2025年视觉感知挑战赛聚焦统一多模态模型,测试其跨任务理解能力。

Perception Test 2025: Challenge Summary and a Unified VQA Extension

  • 设计统一任务接口,要求模型用同一架构处理多种感知任务。
  • 新增多项选择式视频问答等新题型,推动模型通用性评估。
  • 适合关注多模态模型泛化能力的研究者与开发者参考。

第三届视觉感知挑战赛于2025年IEEE/CVF计算机视觉国际会议(ICCV)期间举办,旨在评测当前先进视频模型并衡量多模态感知进展。本届研讨会包含两个特邀赛道:KiVA(图像理解挑战)和Physic-IQ(视频生成挑战)。本文总结了主赛道结果,涵盖既有任务及新增内容。本年度重点强调任务统一性,以更严苛地检验现有SOTA多模态模型。挑战设五个整合赛道:统一视频问答、统一物体与点跟踪、统一动作与声音定位、基于场景的视频问答,以及长达一小时的视频问答,并设有开放投稿的分析与可解释性赛道。其中,统一视频问答引入新子集,将传统感知任务(如点跟踪、时间动作定位)转化为多选题形式的视频问答,使视频-语言模型可直接处理。统一物体与点跟踪合并原物体跟踪与点跟踪任务;统一动作与声音定位则融合时间动作与声音定位任务。参赛者需采用统一方法,而非依赖任务专用模型的流水线方案。该统一挑战凸显当前模型在通过统一接口应对多样化感知任务时的重大挑战。

原文摘要 · Abstract (English)

The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2025. Its primary goal is to benchmark state-of-the-art video models and measure the progress in multimodal perception. This year, the workshop featured 2 guest tracks as well: KiVA (an image understanding challenge) and Physic-IQ (a video generation challenge). In this report, we summarise the results from the main Perception Test challenge, detailing both the existing tasks as well as novel additions to the benchmark. In this iteration, we placed an emphasis on task unification, as this poses a more challenging test for current SOTA multimodal models. The challenge included five consolidated tracks: unified video QA, unified object and point tracking, unified action and sound localisation, grounded video QA, and hour-long video QA, alongside an analysis and interpretability track that is still open for submissions. Notably, the unified video QA track introduced a novel subset that reformulates traditional perception tasks (such as point tracking and temporal action localisation) as multiple-choice video QA questions that video-language models can natively tackle. The unified object and point tracking merged the original object tracking and point tracking tasks, whereas the unified action and sound localisation merged the original temporal action localisation and temporal sound localisation tracks. Accordingly, we required competitors to use unified approaches rather than engineered pipelines with task-specific models. By proposing such a unified challenge, Perception Test 2025 highlights the significant difficulties existing models face when tackling diverse perception tasks through unified interfaces.

多模态视频理解统一建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。