arXiv:2607.05970cs.IRcs.AI2026-07中稿 · SynthIR @ SIGIR 20…

用大模型生成数据集元数据,需权衡搜索效果与真实性

Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search

论文配图:Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
图 1 · 摘自论文原文
  • 对比六种元数据生成方式,从简单重写到基于知识图谱的智能生成
  • 无约束重写提升检索效果但降低真实性,存在语义夸大风险
  • 基于用户画像的生成最平衡,适合追求可信搜索的应用

数据集检索严重依赖元数据,因此大模型生成的元数据成为检索系统中关键的合成内容。我们研究了六种针对RDF数据集的元数据生成方法,涵盖从简单重写到基于用户画像和代理式图结构生成的多种模式,并联合评估其在检索效果与忠实性方面的表现。无约束的元数据重写在检索性能上提升最显著,但忠实性最低,表明搜索优化可能源于未经支持的语义扩展。更受约束的生成方式显著提升了忠实性,其中基于用户画像的重写在检索效果与内容依据之间取得了最佳平衡。这些发现将合成元数据视为一个系统级的信息检索问题,需同时评估效果、来源与可信度。

原文摘要 · Abstract (English)

Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems. We study six metadata-generation settings for RDF datasets, ranging from simple rewriting to profile-grounded and agentic graph-based generation, and evaluate them jointly for retrieval effectiveness and faithfulness. Unconstrained metadata rewriting delivers the strongest retrieval gains over the original metadata, but it is also the least faithful, showing that search improvements can be driven by unsupported semantic expansion. More grounded settings substantially improve faithfulness, and profile-grounded rewriting provides the most balanced trade-off between retrieval effectiveness and grounding. These findings position synthetic metadata as a system-level IR problem in which effectiveness, provenance, and trust must be evaluated together.

元数据生成信息检索大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。