arXiv:2507.01035cs.LGcs.AI2025-07被引 2

优化图神经网络与大模型推荐系统的低延迟推理与训练效率

Research on Low-Latency Inference and Training Efficiency Optimization for Graph Neural Network and Large Language Model-Based Recommendation Systems

  • 融合GNN与LLM,结合量化、LoRA、蒸馏与FPGA加速
  • 最低延迟40-60ms下实现NDCG@10: 0.75,训练时间减少66%
  • 适合追求实时推荐与高效部署的工程团队

在线服务对推荐系统(ReS)的实时性与高效性提出更高要求,需处理复杂的用户-物品交互。本文针对混合图神经网络(GNN)与大语言模型(LLM)的推荐系统,优化其推理延迟与训练效率。采用集成架构优化策略(量化、LoRA、蒸馏)与硬件加速(FPGA、DeepSpeed),在R 4.4.2环境下实验表明:最优配置(Hybrid + FPGA + DeepSpeed)在40-60ms延迟下实现NDCG@10: 0.75,较基线提升13.6%;使用LoRA将训练时间缩短66%(从11.4小时降至3.8小时)。无论在准确性或效率方面,软硬件协同设计与参数高效调优均使混合模型优于独立使用GNN或LLM。建议在实时部署中采用FPGA与LoRA。未来工作可探索联邦学习与先进融合架构以提升可扩展性与隐私保护。本研究为兼顾低延迟与个性化的新一代推荐系统奠定基础。

原文摘要 · Abstract (English)

The incessant advent of online services demands high speed and efficient recommender systems (ReS) that can maintain real-time performance along with processing very complex user-item interactions. The present study, therefore, considers computational bottlenecks involved in hybrid Graph Neural Network (GNN) and Large Language Model (LLM)-based ReS with the aim optimizing their inference latency and training efficiency. An extensive methodology was used: hybrid GNN-LLM integrated architecture-optimization strategies(quantization, LoRA, distillation)-hardware acceleration (FPGA, DeepSpeed)-all under R 4.4.2. Experimental improvements were significant, with the optimal Hybrid + FPGA + DeepSpeed configuration reaching 13.6% more accuracy (NDCG@10: 0.75) at 40-60ms of latency, while LoRA brought down training time by 66% (3.8 hours) in comparison to the non-optimized baseline. Irrespective of domain, such as accuracy or efficiency, it can be established that hardware-software co-design and parameter-efficient tuning permit hybrid models to outperform GNN or LLM approaches implemented independently. It recommends the use of FPGA as well as LoRA for real-time deployment. Future work should involve federated learning along with advanced fusion architectures for better scalability and privacy preservation. Thus, this research marks the fundamental groundwork concerning next-generation ReS balancing low-latency response with cutting-edge personalization.

推荐系统低延迟混合模型参数高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。