提出DRQ方法提升推荐系统中语义ID的鲁棒性
Decoupled Residual Quantization for Robust Semantic IDs in Recommendation
- 分离连续空间重建与离散分布匹配,解耦优化目标
- 量化评估码本重叠与有效容量,诊断失败原因
- 适合关注推荐系统中离散表示质量的研究者
语义ID将物品表示为共享的离散标记序列,已成为推荐与检索的实用工具。然而,难以判断分词器为何失效:质量不佳可能源于码本利用不足、决策边界不稳定或嵌入空间的几何失真。本文提出一个定量分析框架,通过期望码字重叠和有效码本容量来诊断这些缺陷。前者衡量检索扰动下的预期码字混淆,后者将混淆转化为可用且分离良好的码字数量。该框架揭示了语义边界混淆与码字使用不均及欧氏几何约束的关系。作为概念验证,我们提出解耦残差量化(DRQ),将连续几何重建与离散分布匹配分离。在大规模工业数据集上的实验表明,语义ID质量是多目标的:符号鲁棒性、重建保真度和行为感知软匹配分别强调分词器的不同方面。这些下游观察基于单一私有工业数据集,应视为案例研究而非普适基准。
原文摘要 · Abstract (English)
Semantic IDs represent items as shared discrete token sequences and have become a practical tool for recommendation and retrieval. Yet it remains difficult to tell why a tokenizer fails: poor quality may come from codebook underutilization, unstable decision boundaries, or geometric distortion of the embedding space. This paper develops a quantitative framework for diagnosing these failures through expected codeword overlap and effective codebook capacity. The former measures expected codeword confusion under retrieval-time perturbation, while the latter converts that confusion into an effective number of usable, well-separated codes. The framework links semantic boundary confusion to both code usage imbalance and Euclidean geometric constraints. As a proof of concept, we present Decoupled Residual Quantization (DRQ), which separates continuous geometry reconstruction from discrete distribution matching. Experiments on a large-scale industrial dataset show that Semantic ID quality is multi-objective: symbolic robustness, reconstruction fidelity, and behavior-aware soft matching each stress different aspects of a tokenizer. These downstream observations are based on one proprietary industrial dataset, so they should be read as a case study rather than a universal benchmark claim.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。