arXiv:2507.05282cs.IRcs.AI2025-07被引 5

用大模型自动提取数据元数据,提升数据目录构建效率

Exploring LLM Capabilities in Extracting DCAT-Compatible Metadata for Data Cataloging

  • 测试零样本与少样本提示,结合多厂商大模型生成元数据
  • 大模型生成内容质量接近人工,微调后分类准确率显著提升
  • 适合需要快速构建数据目录的机构或团队使用

随着数据在加速流程、改进预测和开发新商业模式中的重要性日益增加,高效的数据探索变得至关重要。由于数据量呈指数增长、异构性强且分布广泛,数据使用者往往花费25%-98%的时间搜寻合适数据。数据目录通过元数据支持和加速数据探索,但元数据的创建与维护通常依赖人工,耗时且需专业知识。本研究探讨大语言模型(LLMs)是否可自动化文本类数据的元数据维护,并生成符合DCAT标准的高质量元数据。我们测试了不同厂商的LLMs在零样本与少样本提示策略下的表现,涵盖标题、关键词生成,以及使用微调模型进行分类。结果表明,大模型在需要高级语义理解的任务中能生成与人工相当的元数据;模型越大性能越好,微调显著提升分类准确率,少样本提示在多数情况下优于零样本。尽管大模型提供了更快、可靠的数据元数据生成方式,但成功应用仍需结合任务特异性标准与领域上下文进行细致考量。

原文摘要 · Abstract (English)

Efficient data exploration is crucial as data becomes increasingly important for accelerating processes, improving forecasts and developing new business models. Data consumers often spend 25-98 % of their time searching for suitable data due to the exponential growth, heterogeneity and distribution of data. Data catalogs can support and accelerate data exploration by using metadata to answer user queries. However, as metadata creation and maintenance is often a manual process, it is time-consuming and requires expertise. This study investigates whether LLMs can automate metadata maintenance of text-based data and generate high-quality DCAT-compatible metadata. We tested zero-shot and few-shot prompting strategies with LLMs from different vendors for generating metadata such as titles and keywords, along with a fine-tuned model for classification. Our results show that LLMs can generate metadata comparable to human-created content, particularly on tasks that require advanced semantic understanding. Larger models outperformed smaller ones, and fine-tuning significantly improves classification accuracy, while few-shot prompting yields better results in most cases. Although LLMs offer a faster and reliable way to create metadata, a successful application requires careful consideration of task-specific criteria and domain context.

大模型数据目录元数据DCAT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。