arXiv:2512.13238cs.CV2025-12被引 3

构建首个真实场景下的双人协作视频语言数据集,用于训练智能助手。

Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance

  • 采用'巫师奥兹'模式收集真人专家指导新手的自然对话。
  • 包含50小时视频与1.5万组高质量视觉问答对,涵盖复杂操作流程。
  • 适合研究人机协作、多模态大模型在真实场景中的应用。

我们提出Ego-EXTRA,一个面向专家-新手协助任务的视频-语言自指视角数据集。该数据集包含50小时未经排练的自指视角视频,记录新手执行程序化任务时,由真实专家通过自然语言提供指导和答疑。采用“巫师奥兹”数据采集范式,专家仅从新手视角观察,主动或被动响应问题,实现双向对话。所有对话被完整录制、转录,形成超过1.5万组高质量视觉问答对,构建全新基准以评估多模态大语言模型。实验表明,该数据集极具挑战性,暴露出当前模型在提供专家级辅助时的显著局限。Ego-EXTRA数据集已公开:https://fpv-iplab.github.io/Ego-EXTRA/。

原文摘要 · Abstract (English)

We present Ego-EXTRA, a video-language Egocentric Dataset for EXpert-TRAinee assistance. Ego-EXTRA features 50 hours of unscripted egocentric videos of subjects performing procedural activities (the trainees) while guided by real-world experts who provide guidance and answer specific questions using natural language. Following a ``Wizard of OZ'' data collection paradigm, the expert enacts a wearable intelligent assistant, looking at the activities performed by the trainee exclusively from their egocentric point of view, answering questions when asked by the trainee, or proactively interacting with suggestions during the procedures. This unique data collection protocol enables Ego-EXTRA to capture a high-quality dialogue in which expert-level feedback is provided to the trainee. Two-way dialogues between experts and trainees are recorded, transcribed, and used to create a novel benchmark comprising more than 15k high-quality Visual Question Answer sets, which we use to evaluate Multimodal Large Language Models. The results show that Ego-EXTRA is challenging and highlight the limitations of current models when used to provide expert-level assistance to the user. The Ego-EXTRA dataset is publicly available to support the benchmark of egocentric video-language assistants: https://fpv-iplab.github.io/Ego-EXTRA/.

视频语言人机协作多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。