评测大模型在真实场景下的跨领域视频理解能力
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering

- 构建四领域第一视角视频问答基准,检验模型泛化性
- 19支队伍胜出开源赛道,38支晋级有限源赛道,超1500份提交
- 资源全公开,适合研究跨域视觉理解与多模态模型的开发者
EgoCross 是一个用于评估多模态大语言模型是否能超越常见日常场景的跨领域第一人称视频问答基准。首届 EgoCross 挑战赛于 CVPR 2026 第三届 EgoVis 工作坊举办,测试模型在四种目标领域(外科手术、工业装配、极限运动、动物视角)的第一人称视频上的表现。每条测试样本包含一段第一人称视频片段、一个问题及四个候选答案,模型需选出正确选项。本技术报告介绍挑战任务、基准资源及两个官方 Codabench 赛道:源受限赛道仅允许使用官方基线模型和少量支持集;开源赛道则允许更广泛模型与训练数据选择,但禁止人工构建目标领域训练数据。挑战共收到来自130余名参与者超过1500份提交,其中19支团队参与开源赛道,38支团队参与源受限赛道。报告公布官方排行榜结果,并总结两赛道优胜方案。所有资源,包括挑战数据、基线实现及获胜团队代码均已公开。
原文摘要 · Abstract (English)
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios. The first EgoCross Challenge was hosted at the Third EgoVis Workshop at CVPR 2026 and evaluated models on first-person videos from four target domains: surgery, industrial assembly, extreme sports, and animal perspectives. Each test example consists of an egocentric video clip, a question, and four candidate answers, from which the model must select the correct option. This technical report introduces the challenge task, benchmark resources, and two official Codabench tracks. The Source-Limited Track restricts participants to the official baseline model and a small support set, whereas the Open-Source Track permits broader choices of models and training data under rules that prohibit the manual construction of target-domain training data. In total, the challenge received more than 1,500 submissions from over 130 participants, with 19 teams participating in the Open-Source Track and 38 teams in the Source-Limited Track. We further present the official leaderboard results and summarize the winning solutions from both tracks. We hope that this report will serve as a useful technical reference for advancing cross-domain egocentric video understanding. All resources, including the challenge data, baseline implementation, and code released by the winning teams, are made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。