用大模型生成描述可显著提升开放数据搜索效果
Keywords are not always the key: A metadata field analysis for natural language search on open data portals
- 分析不同元数据字段对自然语言检索的影响
- 真实数据集测试显示描述字段最关键,大模型生成内容有效
- 适合关注数据可发现性的研究人员与平台开发者
开放数据门户是公众获取开放数据的重要渠道,但其搜索界面通常依赖关键词匹配和有限的元数据字段,难以支持自然语言查询。元数据常不完整或不一致,尤其当用户不熟悉领域术语时问题更严重。本文通过在真实数据集上模拟自然语言查询,开展受控消融实验,评估不同元数据配置下的检索性能,并对比现有‘description’字段与大模型生成内容的质量差异。结果表明,数据集描述在对齐用户意图中起核心作用,且大模型生成的描述能有效提升检索效果。研究揭示了当前元数据实践的局限性,也展示了生成式模型在提升数据可发现性方面的潜力。
原文摘要 · Abstract (English)
Open data portals are essential for providing public access to open datasets. However, their search interfaces typically rely on keyword-based mechanisms and a narrow set of metadata fields. This design makes it difficult for users to find datasets using natural language queries. The problem is worsened by metadata that is often incomplete or inconsistent, especially when users lack familiarity with domain-specific terminology. In this paper, we examine how individual metadata fields affect the success of conversational dataset retrieval and whether LLMs can help bridge the gap between natural queries and structured metadata. We conduct a controlled ablation study using simulated natural language queries over real-world datasets to evaluate retrieval performance under various metadata configurations. We also compare existing content of the metadata field 'description' with LLM-generated content, exploring how different prompting strategies influence quality and impact on search outcomes. Our findings suggest that dataset descriptions play a central role in aligning with user intent, and that LLM-generated descriptions can support effective retrieval. These results highlight both the limitations of current metadata practices and the potential of generative models to improve dataset discoverability in open data portals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。