构建跨领域自拍视频问答基准,检验大模型在真实场景下的泛化能力。
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
- 设计覆盖4类真实场景的跨域评测集,涵盖手术、工业、极限运动等
- 包含约1000个问答对,支持开闭两种问答格式,覆盖四大任务类型
- 揭示现有模型在非日常场景中严重退化,推动鲁棒性研究
多模态大语言模型(MLLMs)在自拍视频问答(EgocentricQA)上取得显著进展,但现有评测主要局限于烹饪、清洁等日常活动。实际应用中常面临视觉风格与语义内容差异大的领域漂移问题。为此,我们提出 extbf{EgoCross},一个全面评估 MLLMs 跨域泛化能力的基准。该数据集涵盖手术、工业、极限运动和动物视角四个多样且具挑战性的领域,代表真实高影响的应用场景。共包含约1,000个问答对,分布在798段视频中,覆盖预测、识别、定位、计数四项关键任务。每对问答提供 OpenQA 与 CloseQA 两种格式,支持细粒度评估。大量实验表明,大多数现有 MLLMs(无论通用或专用于自拍)在超出日常生活的领域上表现不佳,暴露出当前模型的局限性。此外,我们开展了微调与强化学习等初步研究,探索改进路径。我们希望 EgoCross 及配套分析能为发展领域自适应、鲁棒的自拍视频理解奠定基础。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have significantly pushed the frontier of egocentric video question answering (EgocentricQA). However, existing benchmarks and studies are mainly limited to common daily activities such as cooking and cleaning. In contrast, real-world deployment inevitably encounters domain shifts, where target domains differ substantially in both visual style and semantic content. To bridge this gap, we introduce \textbf{EgoCross}, a comprehensive benchmark designed to evaluate the cross-domain generalization of MLLMs in EgocentricQA. EgoCross covers four diverse and challenging domains, including surgery, industry, extreme sports, and animal perspective, representing realistic and high-impact application scenarios. It comprises approximately 1,000 QA pairs across 798 video clips, spanning four key QA tasks: prediction, recognition, localization, and counting. Each QA pair provides both OpenQA and CloseQA formats to support fine-grained evaluation. Extensive experiments show that most existing MLLMs, whether general-purpose or egocentric-specialized, struggle to generalize to domains beyond daily life, highlighting the limitations of current models. Furthermore, we conduct several pilot studies, e.g., fine-tuning and reinforcement learning, to explore potential improvements. We hope EgoCross and our accompanying analysis will serve as a foundation for advancing domain-adaptive, robust egocentric video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。