构建首个面对面教学视频数据集,评测多模态大模型表现
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
- 采集第一人称视角教学互动视频,标注步骤分割与对话状态
- 多模态大模型零样本优于专用模型,展现整体理解潜力
- 适合教育科技、人机交互研究者参考
分析教师与学习者在相同物理空间中的面对面教学互动,是教育支持与技能传递的关键问题,但当前计算机视觉领域尚未系统研究此类场景。我们识别出两大瓶颈:一是缺乏合适的数据集,二是分析方法有限。为此,我们构建了一个新的第一人称视角教学视频数据集,并为两个基础任务提供真实标注:流程步骤分割与对话状态分类。基于该数据集,我们对多模态大语言模型(MLLMs)与传统任务专用模型进行基准测试。由于面对面教学涉及多模态信息(语音内容与语调、视线与身体动作、视觉上下文),有效理解需整合处理言语与非言语交流。因此,我们评估了近期提出的联合处理图像、音频和文本的MLLMs。实验表明,即使未进行任务特定微调,MLLMs也优于专用基线模型,量化了当前机器学习模型对面对面教学场景的理解程度。
原文摘要 · Abstract (English)
Analyzing instructional interactions between an instructor and a learner who are co-present in the same physical space is a critical problem for educational support and skill transfer. Yet such face-to-face instructional scenes have not been systematically studied in computer vision. We identify two key reasons: i) the lack of suitable datasets and ii) limited analytical techniques. To address this gap, we present a new egocentric video dataset of face-to-face instruction and provide ground-truth annotations for two fundamental tasks that serve as a first step toward a comprehensive understanding of instructional interactions: procedural step segmentation and conversation-state classification. Using this dataset, we benchmark multimodal large language models (MLLMs) against conventional task-specific models. Since face-to-face instruction involves multiple modalities (speech content and prosody, gaze and body motion, and visual context), effective understanding requires methods that handle verbal and nonverbal communication in an integrated manner. Accordingly, we evaluate recently introduced MLLMs that jointly process images, audio, and text. This evaluation quantifies the extent to which current machine learning models understand face-to-face instructional scenes. In experiments, MLLMs outperform specialized baselines even without task-specific fine-tuning, suggesting their promise for holistic understanding of instructional interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。