arXiv:2503.09003cs.IRcs.CL2025-03被引 11

用检索增强生成模型自动补全数据目录元信息,提升搜索可用性。

Leveraging Retrieval Augmented Generative LLMs For Automated Metadata Description Generation to Enhance Data Catalogs

  • 结合检索与少样本提示,利用大模型生成高质量元数据描述。
  • 生成内容的Rouge-1 F1超过80%,87%-88%可直接采纳或微调使用。
  • 适合需要规模化优化数据目录的企业数据治理团队。

数据目录是组织中管理与访问数据资产的核心资源,但其有效性依赖于业务用户能否便捷查找相关内容。然而,许多企业数据目录因缺乏完整的元数据(如资产描述)而难以搜索。为此,本文提出一种基于检索增强的少样本提示方法,结合生成式大语言模型(LLM)自动丰富元数据。研究对比了预训练模型(Llama、GPT3.5)与微调后模型(Llama2-7b)在准确性、事实一致性及毒性方面的表现。初步结果显示,生成内容的Rouge-1 F1超过80%,且87%-88%的实例被数据管理员接受为可用或仅需轻微修改。该方法通过自动化生成表与字段描述,构建了一套可扩展的元数据治理框架,显著提升数据目录的可搜索性与整体可用性。

原文摘要 · Abstract (English)

Data catalogs serve as repositories for organizing and accessing diverse collection of data assets, but their effectiveness hinges on the ease with which business users can look-up relevant content. Unfortunately, many data catalogs within organizations suffer from limited searchability due to inadequate metadata like asset descriptions. Hence, there is a need of content generation solution to enrich and curate metadata in a scalable way. This paper explores the challenges associated with metadata creation and proposes a unique prompt enrichment idea of leveraging existing metadata content using retrieval based few-shot technique tied with generative large language models (LLM). The literature also considers finetuning an LLM on existing content and studies the behavior of few-shot pretrained LLM (Llama, GPT3.5) vis-à-vis few-shot finetuned LLM (Llama2-7b) by evaluating their performance based on accuracy, factual grounding, and toxicity. Our preliminary results exhibit more than 80% Rouge-1 F1 for the generated content. This implied 87%- 88% of instances accepted as is or curated with minor edits by data stewards. By automatically generating descriptions for tables and columns in most accurate way, the research attempts to provide an overall framework for enterprises to effectively scale metadata curation and enrich its data catalog thereby vastly improving the data catalog searchability and overall usability.

元数据生成大模型应用数据目录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。