arXiv:2512.02555cs.CLcs.AI2025-12中稿 · SIGIR 2025被引 8

用自动生成数据和认知优化,让电商搜索更懂用户意图。

ADORE: Autonomous Domain-Oriented Relevance Engine for E-commerce

  • 用大模型生成意图对齐的训练数据,再通过行为优化修正。
  • 自动合成对抗样本提升模型鲁棒性,解决数据稀缺问题。
  • 将商品属性结构注入小模型,适合工业部署且推理高效。

电商搜索的相关性建模仍受制于传统词匹配方法(如BM25)的语义鸿沟,以及神经模型对领域内难样本稀缺的依赖。我们提出ADORE,一种自我持续的框架,融合三项创新:(1) 规则感知的相关性判别模块,利用思维链大模型生成意图对齐的训练数据,并通过Kahneman-Tversky优化(KTO)使其与用户行为对齐;(2) 错误类型感知的数据合成模块,自动生成对抗样本以增强模型鲁棒性;(3) 关键属性增强的知识蒸馏模块,将领域特定的属性层级注入可部署的学生模型。ADORE实现了标注、对抗生成与知识蒸馏的自动化,在克服数据稀缺的同时提升了推理能力。大规模实验与线上A/B测试验证了其有效性,为工业应用中的资源高效、认知对齐的相关性建模树立了新范式。

原文摘要 · Abstract (English)

Relevance modeling in e-commerce search remains challenged by semantic gaps in term-matching methods (e.g., BM25) and neural models' reliance on the scarcity of domain-specific hard samples. We propose ADORE, a self-sustaining framework that synergizes three innovations: (1) A Rule-aware Relevance Discrimination module, where a Chain-of-Thought LLM generates intent-aligned training data, refined via Kahneman-Tversky Optimization (KTO) to align with user behavior; (2) An Error-type-aware Data Synthesis module that auto-generates adversarial examples to harden robustness; and (3) A Key-attribute-enhanced Knowledge Distillation module that injects domain-specific attribute hierarchies into a deployable student model. ADORE automates annotation, adversarial generation, and distillation, overcoming data scarcity while enhancing reasoning. Large-scale experiments and online A/B testing verify the effectiveness of ADORE. The framework establishes a new paradigm for resource-efficient, cognitively aligned relevance modeling in industrial applications.

电商搜索大模型应用知识蒸馏自动标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。