用自动生成数据和认知优化,让电商搜索更懂用户意图。
ADORE: Autonomous Domain-Oriented Relevance Engine for E-commerce
- 用大模型生成意图对齐的训练数据,再通过行为优化修正。
- 自动合成对抗样本提升模型鲁棒性,解决数据稀缺问题。
- 将商品属性结构注入小模型,适合工业部署且推理高效。
电商搜索的相关性建模仍受制于传统词匹配方法(如BM25)的语义鸿沟,以及神经模型对领域内难样本稀缺的依赖。我们提出ADORE,一种自我持续的框架,融合三项创新:(1) 规则感知的相关性判别模块,利用思维链大模型生成意图对齐的训练数据,并通过Kahneman-Tversky优化(KTO)使其与用户行为对齐;(2) 错误类型感知的数据合成模块,自动生成对抗样本以增强模型鲁棒性;(3) 关键属性增强的知识蒸馏模块,将领域特定的属性层级注入可部署的学生模型。ADORE实现了标注、对抗生成与知识蒸馏的自动化,在克服数据稀缺的同时提升了推理能力。大规模实验与线上A/B测试验证了其有效性,为工业应用中的资源高效、认知对齐的相关性建模树立了新范式。
原文摘要 · Abstract (English)
Relevance modeling in e-commerce search remains challenged by semantic gaps in term-matching methods (e.g., BM25) and neural models' reliance on the scarcity of domain-specific hard samples. We propose ADORE, a self-sustaining framework that synergizes three innovations: (1) A Rule-aware Relevance Discrimination module, where a Chain-of-Thought LLM generates intent-aligned training data, refined via Kahneman-Tversky Optimization (KTO) to align with user behavior; (2) An Error-type-aware Data Synthesis module that auto-generates adversarial examples to harden robustness; and (3) A Key-attribute-enhanced Knowledge Distillation module that injects domain-specific attribute hierarchies into a deployable student model. ADORE automates annotation, adversarial generation, and distillation, overcoming data scarcity while enhancing reasoning. Large-scale experiments and online A/B testing verify the effectiveness of ADORE. The framework establishes a new paradigm for resource-efficient, cognitively aligned relevance modeling in industrial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。