用大模型让电商定价更智能,兼顾长期收益与可解释性。
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing

- 用大模型结合领域知识与文本上下文生成可解释的定价策略。
- 14天内提升GMV 13.21%、ROI 7.59%、里程碑达成率8.20%。
- 适合关注长期商业目标与决策透明度的电商系统优化者。
大规模电商中的传统动态定价模型存在可解释性差、未结构化信息利用不足,以及与长期业务目标(如累计商品交易额GMV、投资回报率ROI和里程碑达成)不一致的问题。我们提出AIGP框架,通过提示领域知识、结构化数据与文本上下文的大语言模型(LLM),实现可解释且具备知识感知的定价决策。为高效部署并保持高质量输出,采用监督微调进行知识蒸馏。AIGP的核心是离线强化学习训练的长期价值估计算器(LTVE),作为奖励模型对候选定价动作评分,并用于选择偏好对进行直接偏好优化(DPO),从而对齐定价策略与长期业务目标。在淘宝工厂平台的大量离线评估与大规模在线A/B测试表明,相较于生产基线,AIGP在14天内实现GMV提升13.21%、ROI提升7.59%、里程碑达成率提升8.20%,同时提供可解释的定价理由。
原文摘要 · Abstract (English)
Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured information, and misalignment with long-term business objectives such as cumulative Gross Merchandise Value (GMV), Return on Investment (ROI) and milestone achievement. We propose AIGP, a novel framework that leverages a Large Language Model (LLM) prompted with domain knowledge, structured data and textual context to make interpretable, knowledge-aware pricing decisions. For efficient deployment while maintaining high-quality outputs, we employ supervised fine-tuning for knowledge distillation. Central to AIGP is the Long-Term Value Estimator (LTVE), trained via offline reinforcement learning on historical data, which serves as a reward model to score candidate pricing actions and select preference pairs for Direct Preference Optimization (DPO), thereby aligning the pricing policy with long-term business objectives. Extensive offline evaluations and large-scale online A/B tests on Tao Factory demonstrate that AIGP achieves significant improvements: +13.21% in GMV, +7.59% in ROI, and +8.20% in milestone achievement rate over 14 days compared to the production baseline, while simultaneously providing interpretable and transparent pricing rationales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。