构建首个多模态人机协作教学数据集,评估AI能否像真人一样教人用软件。
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching

- 构建72组专家-新手对话数据,含2.8万句对话与28.1小时屏幕记录。
- 模型指令更直接但缺少解释与错误诊断,视觉上下文对齐差。
- 适合研究人机交互、智能辅导系统及教育类AI的开发者参考。
随着智能体在自动化软件任务方面能力增强,它们能否教会人类使用软件?我们提出DigitalCoach,一个包含72个专家-新手计算机使用教学会话的多模态数据集,涵盖22,752轮对话,基于五个软件应用的28.1小时屏幕与输入事件记录。利用该数据集,我们评估了当前最先进的模型在教授用户使用计算机方面的能力。自动评估显示,模型相较于人类在教学方式上存在差异:模型提供更直接的指令,但解释更少、错误诊断更弱、知识检验问题更匮乏。当统一教学方法后,模型生成的话语虽与人类参考语句相似,但在视觉上下文上的对齐性较差。交互式评估进一步证实,模型教练导致学习者被动执行指令,缺乏深层参与,且在视觉信息理解上表现不足。DigitalCoach为协同式与主动式计算机使用辅导智能体的研究奠定了基础。数据与代码已公开于https://project-digital-coach.vercel.app。
原文摘要 · Abstract (English)
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sessions consisting of 22,752 dialogue turns grounded in 28.1 hours of screen and input event recordings across five software applications. We use DigitalCoach to evaluate whether state-of-the-art models can teach humans how to use computers. Automated evaluation shows that models differ from humans in how they coach: models provide more direct instructions, but fewer explanations, error diagnoses, and knowledge-check questions. When we fix the coaching method, models produce utterances similar to human references yet poorly grounded in visual context. Interactive evaluation confirms that model coaches cause learners to passively follow instructions without deeper engagement and fall short in visual grounding. DigitalCoach lays a foundation for collaborative and proactive computer use coaching agents. Data and code are available at https://project-digital-coach.vercel.app.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。