arXiv:2604.03631cs.AI2026-04

用多智能体系统提升屏幕协作学习行为自动分析效果

Single-agent vs. Multi-agents for Automated Video Analysis of On-Screen Collaborative Learning Behaviors

  • 设计两种多智能体框架,分别基于流程化分工和自主决策迭代
  • 多智能体在场景与动作检测上均优于单个大模型,最佳方案准确率超基线
  • 适合教育数据自动化分析、教育技术研究者参考

屏幕学习行为为理解学生信息获取、使用与创造过程提供了重要洞察。分析屏幕行为参与度对捕捉认知与协作过程至关重要。近期视觉语言模型(VLMs)的发展为自动化处理多模态视频数据的繁琐人工编码提供了新机遇。本研究比较了领先闭源VLM(Claude-3.7-Sonnet、GPT-4.1)与开源VLM(Qwen2.5-VL-72B)在单智能体与多智能体设置下的表现,针对协作学习场景中屏幕录制的自动化编码任务,基于ICAP框架进行评估。提出并对比两种多智能体架构:1)三智能体流程式系统,按场景分割视频,结合光标信息提示与证据验证进行行为检测;2)受ReAct启发的自主决策系统,通过推理、工具操作(分割/分类/验证)与观察驱动的自我修正实现可解释标签生成。实验表明,两种多智能体系统均表现出色,在场景与动作检测任务中超越单个VLM。其中流程式系统在场景检测上表现最佳,自主决策系统在动作检测上最优。本研究证明了基于VLM的多智能体系统在视频分析中的有效性,并贡献了一个可扩展的多模态数据分析框架。

原文摘要 · Abstract (English)

On-screen learning behavior provides valuable insights into how students seek, use, and create information during learning. Analyzing on-screen behavioral engagement is essential for capturing students' cognitive and collaborative processes. The recent development of Vision Language Models (VLMs) offers new opportunities to automate the labor-intensive manual coding often required for multimodal video data analysis. In this study, we compared the performance of both leading closed-source VLMs (Claude-3.7-Sonnet, GPT-4.1) and open-source VLM (Qwen2.5-VL-72B) in single- and multi-agent settings for automated coding of screen recordings in collaborative learning contexts based on the ICAP framework. In particular, we proposed and compared two multi-agent frameworks: 1) a three-agent workflow multi-agent system (MAS) that segments screen videos by scene and detects on-screen behaviors using cursor-informed VLM prompting with evidence-based verification; 2) an autonomous-decision MAS inspired by ReAct that iteratively interleaves reasoning, tool-like operations (segmentation/ classification/ validation), and observation-driven self-correction to produce interpretable on-screen behavior labels. Experimental results demonstrated that the two proposed MAS frameworks achieved viable performance, outperforming the single VLMs in scene and action detection tasks. It is worth noting that the workflow-based agent achieved best on scene detection, and the autonomous-decision MAS achieved best on action detection. This study demonstrates the effectiveness of VLM-based Multi-agent System for video analysis and contributes a scalable framework for multimodal data analytics.

视频分析多智能体教育技术VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。