通过自动化数据增强提升银行推荐系统的查询数据质量与规模。
Enhancing and Scaling Search Query Datasets for Recommendation Systems
- 构建三模块系统:生成合成查询、消歧意图、挖掘潜在需求。
- 意图消歧F1达0.863,新意图识别恢复率最高71%。
- 适合需解决冷启动与数据稀疏问题的金融推荐场景。
本文提出一个可部署的生产级系统,用于增强和扩展数字银行业务中基于意图的推荐系统所用的搜索查询数据集。真实环境中用户意图日益复杂多样,导致数据管理困难,影响推荐效果并延迟产品上线。为此,该方法从模型驱动转向自动化数据驱动策略。系统集成三大模块:合成查询生成、意图消歧与意图缺口分析。合成查询生成产生多样化且真实的用户查询;实验显示,对Clinc150使用合成数据无显著差异,而Banking77及专有数据集则存在显著差异,表明该方法有效缓解冷启动问题(即新产品的推荐数据不足)。意图消歧将模糊重叠的意图类别细化为精确子意图,经专家重标注验证,F1得分为0.863 ± 0.127,实现更清晰的推荐映射。意图缺口分析从无标签查询中挖掘潜在客户需求,受控评估下新意图恢复率达71%。系统已在实际银行环境部署,显著提升推荐精度与运营敏捷性,优化用户体验并带来战略收益。本工作强调高质量、可扩展数据在现代AI应用中的关键作用,倡导以主动数据增强驱动价值创造。
原文摘要 · Abstract (English)
This paper presents a deployed, production-grade system designed to enhance and scale search query datasets for intent-based recommendation systems in digital banking. In real-world environments, the growing volume and complexity of user intents create substantial challenges for data management, resulting in suboptimal recommendations and delayed product onboarding. To overcome these challenges, our approach shifts the focus from model-centric enhancements to automated, data-centric strategies. The proposed system integrates three core modules: Synthetic Query Generation, Intent Disambiguation, and Intent Gap Analysis. Synthetic Query Generation produces diverse and realistic user queries. Our experiments reveal no statistically significant difference when using synthetic data for Clinc150, while Banking77 and a proprietary dataset show significant differences. We dig into the underlying factors driving these variations, demonstrating that our approach effectively alleviates the cold start problem (i.e. the challenge of recommending new products with limited historical data). Intent Disambiguation refines broad and overlapping intent categories into precise subintents, achieving an F1 score of 0.863 $\pm$ 0.127 against expert reannotations and leading to clearer differentiation and more precise recommendation mapping. Meanwhile, Intent Gap Analysis identifies latent customer needs by extracting novel intents from unlabeled queries; recovery rates reach up to 71\% in controlled evaluations. Deployed in a live banking environment, our system demonstrates significant improvements in recommendation precision and operation agility, ultimately delivering enhanced user experiences and strategic business benefits. This work underscores the role of high-quality, scalable data in modern AI-driven applications and advocates a proactive approach to data enhancement as a key driver of value.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。