arXiv:2505.07245cs.LGcs.AI2025-05

解决购车预测中的极端不平衡问题,提升精准推荐效率。

REMEDI: Relative Feature Enhanced Meta-Learning with Distillation for Imbalanced Prediction

  • 用多模型融合与相对性能特征增强元学习。
  • 在80万用户上实现前6万推荐中10%精度,覆盖50%真实买家。
  • 适合工业场景下高精度、低延迟的不平衡预测任务。

现有车主未来购车预测面临极端类别不平衡(正样本率低于0.5%)和复杂行为模式的挑战。本文提出REMEDI(基于蒸馏的相对特征增强元学习框架),分三阶段应对:首先训练多个基础模型以捕捉用户行为的互补特征;其次借鉴比较优化思想,引入相对性能元特征(偏离集成均值、同侪排名),通过混合专家架构实现高效模型融合;最后采用带MSE损失的监督微调,将集成知识蒸馏至单个高效模型,便于实际部署。在约80万车主数据集上评估,REMEDI显著优于基线方法,在前6万推荐中实现约10%精度,覆盖约50%的真实购买者,且蒸馏模型保持了集成性能的同时具备部署效率,验证了其在工业级不平衡预测中的有效性。

原文摘要 · Abstract (English)

Predicting future vehicle purchases among existing owners presents a critical challenge due to extreme class imbalance (<0.5% positive rate) and complex behavioral patterns. We propose REMEDI (Relative feature Enhanced Meta-learning with Distillation for Imbalanced prediction), a novel multi-stage framework addressing these challenges. REMEDI first trains diverse base models to capture complementary aspects of user behavior. Second, inspired by comparative op-timization techniques, we introduce relative performance meta-features (deviation from ensemble mean, rank among peers) for effective model fusion through a hybrid-expert architecture. Third, we distill the ensemble's knowledge into a single efficient model via supervised fine-tuning with MSE loss, enabling practical deployment. Evaluated on approximately 800,000 vehicle owners, REMEDI significantly outperforms baseline approaches, achieving the business target of identifying ~50% of actual buyers within the top 60,000 recommendations at ~10% precision. The distilled model preserves the ensemble's predictive power while maintaining deployment efficiency, demonstrating REMEDI's effectiveness for imbalanced prediction in industry settings.

不平衡学习元学习模型蒸馏推荐系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。