arXiv:2601.13876cs.CL2026-01被引 1

让轻量机器人模型讲出有教育意义的科学实验解说。

Pedagogical Alignment for Vision-Language-Action Models: A Comprehensive Framework for Data, Architecture, and Evaluation in Education

  • 用文本修复和知识蒸馏恢复小模型的语言生成能力
  • 在五类科学实验中任务成功率与基线相当,且能生成合适解释
  • 适合资源受限的课堂,尤其需要可解释讲解的教育场景

科学演示对有效开展科学、技术、工程和数学教育至关重要,但教师在多次教学中安全且一致地实施演示面临挑战,此时机器人可提供帮助。然而,现有视觉-语言-动作(VLA)模型需大量计算资源,并为提升效率牺牲语言生成能力,难以适用于需可解释性与生成讲解的资源受限教育环境。本文提出「教育对齐VLA框架」,通过四个组件实现轻量VLA模型的教育对齐:文本修复以恢复语言生成能力,大语言模型蒸馏以传递教育知识,面向教育环境的安全训练,以及针对科学教育情境调整的教育评估。我们在物理、化学、生物和地球科学共五项科学演示中评估该框架,评估体系由科学教育专家共同设计。评估涵盖任务表现(成功率、流程合规性、效率、安全性)及教育质量,通过教师问卷和大语言模型作为裁判进行评分。结果表明,该框架在任务表现上达到基线水平的同时,能够生成符合上下文的教育性解释,并通过生成文本的定性分析验证其合理性。

原文摘要 · Abstract (English)

Science demonstrations are important for effective STEM education, yet teachers face challenges in conducting them safely and consistently across multiple occasions, where robotics can be helpful. However, current Vision-Language-Action (VLA) models require substantial computational resources and sacrifice language generation capabilities to maximize efficiency, making them unsuitable for resource-constrained educational settings that require interpretable, explanation-generating systems. We present \textit{Pedagogical VLA Framework}, a framework that applies pedagogical alignment to lightweight VLA models through four components: text healing to restore language generation capabilities, large language model (LLM) distillation to transfer pedagogical knowledge, safety training for educational environments, and pedagogical evaluation adjusted to science education contexts. We evaluate Pedagogical VLA Framework across five science demonstrations spanning physics, chemistry, biology, and earth science, using an evaluation framework developed in collaboration with science education experts. Our evaluation assesses both task performance (success rate, protocol compliance, efficiency, safety) and pedagogical quality through teacher surveys and LLM-as-Judge assessment. We additionally provide qualitative analysis of generated texts. Experimental results demonstrate that Pedagogical VLA Framework achieves comparable task performance to baseline models while producing contextually appropriate educational explanations.

教育AIVLA模型轻量化科学教学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。