用共享权重自动设计推荐模型,兼顾架构与硬件协同优化。
Towards Automated Model Design on Recommender Systems
- 构建超网络统一搜索模型架构与硬件配置
- 在三个点击率预测任务中超越现有模型性能
- 实现2倍算力效率提升,适合系统级优化研究者
深度学习的普及为基于AI的推荐系统带来新机遇。使用深度神经网络设计推荐系统需精细架构设计,而进一步优化则需联合模型架构与硬件进行协同设计。自动化方法(如AutoML)对于充分挖掘推荐模型设计潜力至关重要,包括模型选择与模型-硬件协同设计策略。本文提出一种新范式,利用权重共享探索海量解空间。通过构建大型超网络,搜索最优架构与协同设计策略,以应对推荐领域中的多模态与异构数据挑战。从模型角度,超网络包含多种操作符、密集连接与维度搜索选项;从协同设计角度,涵盖多样化的存算一体(PIM)配置,生成硬件高效模型。其规模、异构性与复杂性带来多重挑战,我们提出一系列训练与评估技术予以解决。所设计模型在三个点击率(CTR)预测基准上表现优异,优于人工设计及现有AutoML方法,达到当前最优架构搜索性能。从协同设计角度看,实现2倍浮点运算效率提升、1.8倍能效提升和1.5倍性能提升。
原文摘要 · Abstract (English)
The increasing popularity of deep learning models has created new opportunities for developing AI-based recommender systems. Designing recommender systems using deep neural networks requires careful architecture design, and further optimization demands extensive co-design efforts on jointly optimizing model architecture and hardware. Design automation, such as Automated Machine Learning (AutoML), is necessary to fully exploit the potential of recommender model design, including model choices and model-hardware co-design strategies. We introduce a novel paradigm that utilizes weight sharing to explore abundant solution spaces. Our paradigm creates a large supernet to search for optimal architectures and co-design strategies to address the challenges of data multi-modality and heterogeneity in the recommendation domain. From a model perspective, the supernet includes a variety of operators, dense connectivity, and dimension search options. From a co-design perspective, it encompasses versatile Processing-In-Memory (PIM) configurations to produce hardware-efficient models. Our solution space's scale, heterogeneity, and complexity pose several challenges, which we address by proposing various techniques for training and evaluating the supernet. Our crafted models show promising results on three Click-Through Rates (CTR) prediction benchmarks, outperforming both manually designed and AutoML-crafted models with state-of-the-art performance when focusing solely on architecture search. From a co-design perspective, we achieve 2x FLOPs efficiency, 1.8x energy efficiency, and 1.5x performance improvements in recommender models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。