首个面向物理实验的多粒度视觉解析数据集,助力智能教育发展。
PhysLab: A Benchmark Dataset for Multi-Granularity Visual Parsing of Physics Experiments
- 构建学生做物理实验的长视频数据集,支持细粒度动作与交互分析。
- 含620段视频、4类实验,覆盖多样仪器与人-物交互模式。
- 适用于教育场景视觉理解,推动计算机视觉与教学技术融合。
图像与视频的视觉解析在众多现实应用中至关重要。然而当前研究受限于现有数据集的不足:(1)标注粒度不够,难以支持细粒度场景理解与高层推理;(2)领域覆盖有限,尤其缺乏面向教育场景的数据集;(3)缺乏明确的流程指引,逻辑规则少,任务过程结构化表示不足。为此,我们提出PhysLab,首个记录学生完成复杂物理实验的视频数据集。该数据集包含4个代表性实验,涵盖多种科学仪器与丰富的人员-物体交互(HOI)模式。PhysLab共包含620段长视频,提供多层级标注,支持动作识别、目标检测、HOI分析等多种视觉任务。我们建立了强基线并进行了广泛评估,揭示了程序性教育视频解析的关键挑战。预期PhysLab将成为推进细粒度视觉解析、促进智能课堂系统发展的宝贵资源。数据集与评估工具已公开于https://github.com/ZMH-SDUST/PhysLab。
原文摘要 · Abstract (English)
Visual parsing of images and videos is critical for a wide range of real-world applications. However, progress in this field is constrained by limitations of existing datasets: (1) insufficient annotation granularity, which impedes fine-grained scene understanding and high-level reasoning; (2) limited coverage of domains, particularly a lack of datasets tailored for educational scenarios; and (3) lack of explicit procedural guidance, with minimal logical rules and insufficient representation of structured task process. To address these gaps, we introduce PhysLab, the first video dataset that captures students conducting complex physics experiments. The dataset includes four representative experiments that feature diverse scientific instruments and rich human-object interaction (HOI) patterns. PhysLab comprises 620 long-form videos and provides multilevel annotations that support a variety of vision tasks, including action recognition, object detection, HOI analysis, etc. We establish strong baselines and perform extensive evaluations to highlight key challenges in the parsing of procedural educational videos. We expect PhysLab to serve as a valuable resource for advancing fine-grained visual parsing, facilitating intelligent classroom systems, and fostering closer integration between computer vision and educational technologies. The dataset and the evaluation toolkit are publicly available at https://github.com/ZMH-SDUST/PhysLab.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。