arXiv:2410.04231cs.IR2024-10被引 7

用检索增强生成提升多源数据发现效率

Metadata-based Data Exploration with Retrieval-Augmented Generation for Large Language Models

  • 结合LLM与向量库,挖掘异构数据间的语义关联
  • 在四类任务中显著提升跨类别数据匹配效果
  • 适合需要跨领域数据探索的研究者使用

面对可用元数据有限的挑战,有效搜索所需数据集成为迫切需求。本研究提出一种基于检索增强生成(RAG)的新架构,用于增强基于元数据的数据发现能力。系统将大语言模型(LLMs)与外部向量数据库结合,识别不同数据类型之间的语义关系,提供一种评估异构数据源间语义相似性的新方法,并改进数据探索流程。实验涵盖四项关键任务:1)推荐相似数据集,2)建议可组合数据集,3)估计标签,4)预测变量。结果表明,相比传统元数据方法,RAG能显著提升跨类别数据的选择准确性,但各项任务表现随模型和任务而异,凸显需根据具体场景选择技术。研究验证了该方法在数据发现中的潜力,但估计类任务仍需进一步优化。

原文摘要 · Abstract (English)

Developing the capacity to effectively search for requisite datasets is an urgent requirement to assist data users in identifying relevant datasets considering the very limited available metadata. For this challenge, the utilization of third-party data is emerging as a valuable source for improvement. Our research introduces a new architecture for data exploration which employs a form of Retrieval-Augmented Generation (RAG) to enhance metadata-based data discovery. The system integrates large language models (LLMs) with external vector databases to identify semantic relationships among diverse types of datasets. The proposed framework offers a new method for evaluating semantic similarity among heterogeneous data sources and for improving data exploration. Our study includes experimental results on four critical tasks: 1) recommending similar datasets, 2) suggesting combinable datasets, 3) estimating tags, and 4) predicting variables. Our results demonstrate that RAG can enhance the selection of relevant datasets, particularly from different categories, when compared to conventional metadata approaches. However, performance varied across tasks and models, which confirms the significance of selecting appropriate techniques based on specific use cases. The findings suggest that this approach holds promise for addressing challenges in data exploration and discovery, although further refinement is necessary for estimation tasks.

数据发现检索增强大模型应用语义匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。