arXiv:2509.11947cs.CYcs.AI2025-09

用量化模型+GPU加速,打造可私有部署的并行计算辅导机器人。

A GPU-Accelerated RAG-Based Telegram Assistant for Supporting Parallel Processing Students

  • 基于量化Mistral-7B Instruct与RAG,实现课程知识精准问答
  • 在消费级GPU上实现低延迟推理,支持实时响应
  • 适合需随时获取并行计算辅导的学生和教育者

本项目解决学生在常规授课时间外持续获得学术支持的关键需求。提出一个面向特定领域‘并行处理导论’课程的检索增强生成(RAG)系统,采用量化版Mistral-7B Instruct模型,并部署为Telegram聊天机器人。该助手通过实时、个性化的回复,帮助学生理解课程内容。利用GPU加速显著降低推理延迟,使系统可在消费级硬件上实用化部署。结果表明,消费级GPU即可实现低成本、私密且高效的高性能计算教育AI辅导。

原文摘要 · Abstract (English)

This project addresses a critical pedagogical need: offering students continuous, on-demand academic assistance beyond conventional reception hours. I present a domain-specific Retrieval-Augmented Generation (RAG) system powered by a quantized Mistral-7B Instruct model and deployed as a Telegram bot. The assistant enhances learning by delivering real-time, personalized responses aligned with the "Introduction to Parallel Processing" course materials. GPU acceleration significantly improves inference latency, enabling practical deployment on consumer hardware. This approach demonstrates how consumer GPUs can enable affordable, private, and effective AI tutoring for HPC education.

RAGAI辅导并行计算量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。