探究编程初学者如何选择多模态AI工具及决策依据。
Exploring Student Choice and the Use of Multimodal Generative AI in Programming Learning
- 通过16次思考录音实验,观察学生使用多模态AI解题时的模态选择行为。
- 发现学生根据问题类型和自身理解方式灵活切换文本、图像、语音等交互模式。
- 为未来教育类多模态AI设计提供实证依据,适合教育技术研究者参考。
生成式人工智能(GenAI)的广泛应用正在影响计算机科学教育,已有研究揭示其在编程学习中带来的益处与潜在风险。然而,多数现有研究聚焦于仅支持文本交互的GenAI工具。随着技术发展,GenAI已开始支持多种交互模式,即多模态交互。本文探讨了本科生编程初学者在使用多模态GenAI工具时的选择行为及其决策标准。研究选用一款支持文本、音频、图像上传和实时屏幕共享的商用多模态GenAI平台,通过16次结合思考录音与后续半结构化访谈的参与者观察实验,分析学生在完成编程任务时对不同交互模态的选择及其背后原因。随着多模态通信成为教育AI的未来方向,本研究旨在推动对计算机科学教育中学生与多模态GenAI互动机制的持续探索。
原文摘要 · Abstract (English)
The broad adoption of Generative AI (GenAI) is impacting Computer Science education, and recent studies found its benefits and potential concerns when students use it for programming learning. However, most existing explorations focus on GenAI tools that primarily support text-to-text interaction. With recent developments, GenAI applications have begun supporting multiple modes of communication, known as multimodality. In this work, we explored how undergraduate programming novices choose and work with multimodal GenAI tools, and their criteria for choices. We selected a commercially available multimodal GenAI platform for interaction, as it supports multiple input and output modalities, including text, audio, image upload, and real-time screen-sharing. Through 16 think-aloud sessions that combined participant observation with follow-up semi-structured interviews, we investigated student modality choices for GenAI tools when completing programming problems and the underlying criteria for modality selections. With multimodal communication emerging as the future of AI in education, this work aims to spark continued exploration on understanding student interaction with multimodal GenAI in the context of CS education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。