arXiv:2502.17060cs.LGstat.ML2025-02被引 1

用向量嵌入统一建模多数据集,提升分析效率与准确性

Analytics Modelling over Multiple Datasets using Vector Embeddings

  • 通过深度模型NumTabData2Vec将数据集转为向量嵌入,实现相似性搜索
  • 相比现有框架,预测准确率更高,执行速度提升显著
  • 可精准映射并区分真实场景,适合数据选型与分析优化

随着数据量和数据集数量的激增,分析师面临如何选择高质量数据以提升分析性能的挑战。为解决这一问题,我们提出一种新方法:利用可用数据集构建模型,预测分析操作的结果。每个数据集通过我们提出的深度学习模型NumTabData2Vec转换为向量嵌入表示,并采用相似性搜索。实验表明,该框架在预测性能和执行时间上均优于当前最先进的建模框架,能更准确地预测分析结果并实现显著加速。此外,该向量化模型可将不同真实场景映射到低维向量空间,并保持良好区分能力。

原文摘要 · Abstract (English)

The massive increase in the data volume and dataset availability for analysts compels researchers to focus on data content and select high-quality datasets to enhance the performance of analytics operators. While selecting high-quality data significantly boosts analytical accuracy and efficiency, the exact process is very challenging given large-scale dataset availability. To address this issue, we propose a novel methodology that infers the outcome of analytics operators by creating a model from the available datasets. Each dataset is transformed to a vector embedding representation generated by our proposed deep learning model NumTabData2Vec, where similarity search are employed. Through experimental evaluation, we compare the prediction performance and the execution time of our framework to another state-of-the-art modelling operator framework, illustrating that our approach predicts analytics outcomes accurately, and increases speedup. Furthermore, our vectorization model can project different real-world scenarios to a lower vector embedding representation accurately and distinguish them.

数据选型向量嵌入分析优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。