arXiv:2410.20325cs.LGcs.SI2024-10

用结构化数据过滤噪声,提升领域特定嵌入的精准度。

Domain Specific Data Distillation and Multi-modal Embedding Generation

  • 基于混合协同过滤框架,用结构化数据优化无结构数据
  • 在云领域任务中精度提升28%,召回率提高11%
  • 适合需要高精度属性预测的领域应用

领域特定嵌入的构建面临非结构化数据泛滥与领域结构化数据稀缺的矛盾。传统嵌入方法通常仅依赖单一模态,适用性受限。本文提出一种新方法,利用结构化数据从非结构化数据中过滤噪声,生成高精度、高召回率的领域属性预测嵌入。模型基于混合协同过滤(HCF)框架,在通用实体表示基础上通过相关项目预测任务进行微调。在云计算领域的实验表明,基于HCF的嵌入优于纯非结构化数据训练的自编码器嵌入,实现领域属性预测上28%的精度提升和11%的召回率提升。

原文摘要 · Abstract (English)

The challenge of creating domain-centric embeddings arises from the abundance of unstructured data and the scarcity of domain-specific structured data. Conventional embedding techniques often rely on either modality, limiting their applicability and efficacy. This paper introduces a novel modeling approach that leverages structured data to filter noise from unstructured data, resulting in embeddings with high precision and recall for domain-specific attribute prediction. The proposed model operates within a Hybrid Collaborative Filtering (HCF) framework, where generic entity representations are fine-tuned through relevant item prediction tasks. Our experiments, focusing on the cloud computing domain, demonstrate that HCF-based embeddings outperform AutoEncoder-based embeddings (using purely unstructured data), achieving a 28% lift in precision and an 11% lift in recall for domain-specific attribute prediction.

嵌入学习协同过滤云平台数据蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。