通过细粒度相关性控制生成合成数据,提升短视频搜索排序效果
Semi-Supervised Synthetic Data Generation with Fine-Grained Relevance Control for Short Video Search Relevance Modeling
- 半监督框架协同生成适应领域的短视频数据,可调相关性标签
- 合成数据覆盖低频中间相关性层级,使训练集更均衡丰富
- 线上A/B测试显示点击率提升1.45%,强相关比例上升4.9%
合成数据广泛用于嵌入模型以增强训练数据在难度、长度和语言等维度的多样性。然而,现有基于提示的合成方法难以捕捉领域特定的数据分布,尤其在数据稀缺领域,且常忽略细粒度相关性多样性。本文构建了一个包含四层级相关性标注的中文短视频数据集,填补关键资源空白。进一步提出一种半监督合成数据流程,两个协同训练的模型可生成具有可控相关性标签的领域自适应短视频数据。该方法通过合成代表性不足的中间相关性样本,提升了相关性层级的多样性,使训练数据更平衡且语义丰富。大量离线实验表明,基于合成数据训练的嵌入模型优于基于提示生成或纯监督微调的数据。此外,实验验证在训练中引入更多细粒度相关性层级能增强模型对细微语义差异的敏感性,凸显细粒度相关性监督在嵌入学习中的价值。在抖音双列推荐场景的搜索增强推荐链路中,线上A/B测试显示,所提模型将点击率(CTR)提升1.45%,强相关比例(SRR)提高4.9%,图像用户渗透率(IUPR)提升0.1054%。
原文摘要 · Abstract (English)
Synthetic data is widely adopted in embedding models to ensure diversity in training data distributions across dimensions such as difficulty, length, and language. However, existing prompt-based synthesis methods struggle to capture domain-specific data distributions, particularly in data-scarce domains, and often overlook fine-grained relevance diversity. In this paper, we present a Chinese short video dataset with 4-level relevance annotations, filling a critical resource void. Further, we propose a semi-supervised synthetic data pipeline where two collaboratively trained models generate domain-adaptive short video data with controllable relevance labels. Our method enhances relevance-level diversity by synthesizing samples for underrepresented intermediate relevance labels, resulting in a more balanced and semantically rich training data set. Extensive offline experiments show that the embedding model trained on our synthesized data outperforms those using data generated based on prompting or vanilla supervised fine-tuning(SFT). Moreover, we demonstrate that incorporating more diverse fine-grained relevance levels in training data enhances the model's sensitivity to subtle semantic distinctions, highlighting the value of fine-grained relevance supervision in embedding learning. In the search enhanced recommendation pipeline of Douyin's dual-column scenario, through online A/B testing, the proposed model increased click-through rate(CTR) by 1.45%, raised the proportion of Strong Relevance Ratio (SRR) by 4.9%, and improved the Image User Penetration Rate (IUPR) by 0.1054%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。