arXiv:2411.04341cs.LG2024-11被引 8

用RAG+课程资料提升中小学教学,验证数据源有效性

Enhancing classroom teaching with LLMs and RAG

  • 以课程材料为知识库,构建RAG教学辅助系统
  • 实验显示不同分块大小下答案准确率均未超50%
  • 方法可评估教育数据源质量,适合教育AI研究者

大型语言模型虽能提供日常信息,但训练后数据易过时,RAG可补充最新内容。本文研究将课程材料作为数据源,构建RAG系统以支持K-12教育。初步实验使用Reddit作为实时网络安全信息源,测试不同分块大小对回答准确率的影响。通过RAGAs评估,所有分块大小下的平均答案正确率均未超过50%,表明Reddit不适合作为网络安全威胁问题的数据来源。该方法成功验证了数据源的可靠性,对评估教育类资源的有效性具有重要启示。

原文摘要 · Abstract (English)

Large Language Models have become a valuable source of information for our daily inquiries. However, after training, its data source quickly becomes out-of-date, making RAG a useful tool for providing even more recent or pertinent data. In this work, we investigate how RAG pipelines, with the course materials serving as a data source, might help students in K-12 education. The initial research utilizes Reddit as a data source for up-to-date cybersecurity information. Chunk size is evaluated to determine the optimal amount of context needed to generate accurate answers. After running the experiment for different chunk sizes, answer correctness was evaluated using RAGAs with average answer correctness not exceeding 50 percent for any chunk size. This suggests that Reddit is not a good source to mine for data for questions about cybersecurity threats. The methodology was successful in evaluating the data source, which has implications for its use to evaluate educational resources for effectiveness.

教育AIRAG教学增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。