arXiv:2609.04550cs.CV2026-09

用视觉语言模型分析课堂视频,实现24类标签的高精度密集标注。

VISTA: Dense Multi-Label Classroom Coding with Vision-Language Models

  • 基于COPUS协议构建视频基准,每2分钟生成24维二值标签
  • 新模型VISTA在化学课上达80.1%精确率,优于零样本74.9%
  • 适合教育技术与多模态模型研究者使用

视频-语言基准通常由数据集作者构建,缺乏公开的可靠性统计,导致标注噪声水平未知。我们主张借鉴已有可靠性保障体系的领域经验来提升多模态基准质量。以本科STEM课堂观察协议(COPUS)为例,该协议包含24个类别标签,已有十年同行评审的可靠性研究。我们将COPUS重构为多模态基础模型的视频基准,提供密集结构化标签(每2分钟一个24维二值向量,覆盖50-90分钟讲座)、经外部验证的词汇表及基于人类评估者的每类可靠性目标。评估语料的标注由5人评审小组达成共识作为参考标准。我们提出VISTA:在密集滑动窗口上运行MiniCPM-V-4.5,通过轻量级MLP头对每窗口输出进行微调,并将结果最大池化至2分钟的COPUS网格。在三段独立化学课上,VISTA达到80.1%受限宏平均准确率,优于零样本版本的74.9%;主要残差错误出现在视觉相似的教师行为标签和依赖音频的罕见标签上。我们识别出三种系统性失败模式(音频部分可观测性、细粒度小组协作区分、长尾召回),并发布基准工具链、提示模板和基线代码于https://github.com/ajfranck/VISTA。

原文摘要 · Abstract (English)

Video-language benchmarks are usually constructed by the dataset authors without published reliability statistics, leaving the noise floor of the construct unknown. We argue that multimodal benchmarking benefits from methods taken from research communities that have already invested in strategies to ensure reliability. We illustrate the case with the Classroom Observation Protocol for Undergraduate STEM (COPUS): a 24-code multi-label observation instrument with a decade of peer-reviewed reliability literature. We recast COPUS as a video benchmark for multimodal foundation models, where it provides a dense set of structured labels (a 24-dimensional binary vector every 2 minutes across a 50-90 minute lecture), an externally validated vocabulary, and established literature that provides a per-code reliability target based on human evaluators. Annotations in our evaluation corpus are produced by a 5-person human-evaluator panel whose consensus matrix is our reference. We propose VISTA, a baseline that runs MiniCPM-V-4.5 over a dense sliding window, refines its per-window outputs with a lightweight multi-layer perceptron (MLP) head trained on top of the frozen backbone, and max-pools the resulting predictions onto the 2-minute COPUS grid. On three held-out chemistry lectures, VISTA reaches 80.1% restricted macro accuracy versus 74.9% for the zero-shot variant, with the largest residual errors on visually similar instructor codes and on rare audio-dependent codes. We characterize three systematic failure modes (audio-partial observability, fine-grained group-work discrimination, long-tail recall) and release the benchmark tooling, prompts and baseline code at https://github.com/ajfranck/VISTA.

多模态课堂分析视觉语言模型教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。