解决电商大模型幻觉与安全漏洞,提升搜索精准与可信度。
SIA: A Synthesize-Inject-Align Framework for Knowledge-Grounded and Secure E-commerce Search LLMs with Industrial Deployment
- 合成知识与行为数据,构建高质量训练语料。
- 参数高效预训练使模型兼顾领域知识与通用能力。
- 双路径对齐提升搜索效果与抗攻击能力,适合工业级部署。
大语言模型为电商搜索带来意图感知推荐的潜力,但其工业应用受限于两大挑战:一是动态细粒度商品知识编码不足导致的知识幻觉,二是越狱攻击引发的安全漏洞威胁合规性。为此,我们提出SIA——一种合成-注入-对齐框架,用于构建具备知识与安全性的电商搜索大模型。首先,通过融合结构化知识图谱与非结构化行为日志,结合推理链和安全增强数据,合成高质量自然语言语料。其次,采用基于深度上扩的参数高效预训练策略,在注入领域知识的同时保持通用能力。最后,通过多任务指令微调与对抗训练的双路径对齐方法,强化任务性能与安全鲁棒性。该框架已在京东(中国最大自营电商平台)部署,五类核心搜索场景的A/B测试显示关键业务指标显著提升,验证了其工业有效性与可扩展性。
原文摘要 · Abstract (English)
Large language models offer transformative potential for e-commerce search by enabling intent-aware recommendations. However, their industrial deployment is hindered by two critical challenges: (1) knowledge hallucination due to insufficient encoding of dynamic, fine-grained product knowledge, and (2) security vulnerabilities under jailbreak attacks that threaten compliance. To address these issues, we propose SIA--a Synthesize-Inject-Align framework for building knowledgeable and secure e-commerce search LLMs. Our approach first synthesizes high-quality natural language corpus by combining structured knowledge graphs with unstructured behavioral logs, augmented with reasoning chains and safety-aware data. We then introduce a parameter-efficient pre-training strategy based on Depth Up-Scaling to inject domain knowledge while preserving general capabilities. Finally, a dual-path alignment method via multi-task instruction tuning and adversarial training strengthens both task performance and safety robustness. The framework has been deployed at JD.com, China's largest self-operated e-commerce platform, where A/B tests across five core search scenarios demonstrate significant improvements in key business metrics, validating its industrial effectiveness and scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。