测试大模型在计算机教育中的表现,发现它查资料强但分析弱。
Can LLMs Assist Computer Education? an Empirical Case Study of DeepSeek
- 用真实网络题和模拟题评估模型表现
- 查事实准确率高,复杂推理能力不足
- 适合辅助教学,不适用于深度分析
本研究通过实证案例评估了新兴大语言模型 DeepSeek-V3 在计算机教育中的有效性与可靠性。评测涵盖 CCNA 模拟题和中国网络工程师提出的实际网络安全部署问题。评估维度包括角色依赖性、跨语言能力及答案可复现性,并进行统计分析。结果表明,无论提示中是否包含角色定义,模型表现均保持稳定;在原始数据与翻译后数据上,准确率也保持一致,证实其跨语言适应能力。模型在低阶事实记忆任务中表现优异,但在高阶推理任务中明显下降,凸显其信息检索优势与复杂分析局限。尽管在网络安全教育中具备显著实用价值,但其处理多模态数据及深入复杂议题的能力仍受限。研究为专业领域大模型的优化提供了重要参考。
原文摘要 · Abstract (English)
This study presents an empirical case study to assess the efficacy and reliability of DeepSeek-V3, an emerging large language model, within the context of computer education. The evaluation employs both CCNA simulation questions and real-world inquiries concerning computer network security posed by Chinese network engineers. To ensure a thorough evaluation, diverse dimensions are considered, encompassing role dependency, cross-linguistic proficiency, and answer reproducibility, accompanied by statistical analysis. The findings demonstrate that the model performs consistently, regardless of whether prompts include a role definition or not. In addition, its adaptability across languages is confirmed by maintaining stable accuracy in both original and translated datasets. A distinct contrast emerges between its performance on lower-order factual recall tasks and higher-order reasoning exercises, which underscores its strengths in retrieving information and its limitations in complex analytical tasks. Although DeepSeek-V3 offers considerable practical value for network security education, challenges remain in its capability to process multimodal data and address highly intricate topics. These results provide valuable insights for future refinement of large language models in specialized professional environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。