用大模型语义搜索提升社科数据发现效率,尤其擅长处理错别字和复杂查询。
Comparing how Large Language Models perform against keyword-based searches for social science research data discovery
- 用语义搜索替代关键词匹配,理解自然语言意图。
- 语义搜索结果数量多30%以上,对模糊查询召回率提升显著。
- 适合研究人员快速探索未知数据,弥补传统搜索盲区。
本文评估了基于大语言模型(LLM)的语义搜索工具在社科研究数据发现中的表现,相较于传统的关键词搜索。基于2023年12月至2024年10月期间从消费数据研究中心(CDRC)搜索日志中提取的131个高频搜索词,比较了定制化语义搜索系统与CDRC关键词搜索的结果。通过描述性统计、定性分析及量化相似度指标(包括精确数据集重叠率、杰卡德相似度、基于BERT嵌入的余弦相似度)评估返回数据集的数量、重合度、排序与相关性。结果显示,语义搜索持续返回更多结果,尤其在地理类、拼写错误、生僻或复杂查询上表现优异。尽管未覆盖所有关键词搜索结果,但返回数据集在语义上高度相似,余弦相似度高而精确重合度较低。两者最相关结果的排序差异显著,反映不同的优先策略。案例研究显示,该工具对拼写错误具有鲁棒性,能有效理解地理与上下文相关性,并支持关键词搜索无法处理的自然语言查询。总体而言,基于大模型的语义搜索在数据发现中表现出显著优势,应作为传统关键词搜索的补充而非替代。
原文摘要 · Abstract (English)
This paper evaluates the performance of a large language model (LLM) based semantic search tool relative to a traditional keyword-based search for data discovery. Using real-world search behaviour, we compare outputs from a bespoke semantic search system applied to UKRI data services with the Consumer Data Research Centre (CDRC) keyword search. Analysis is based on 131 of the most frequently used search terms extracted from CDRC search logs between December 2023 and October 2024. We assess differences in the volume, overlap, ranking, and relevance of returned datasets using descriptive statistics, qualitative inspection, and quantitative similarity measures, including exact dataset overlap, Jaccard similarity, and cosine similarity derived from BERT embeddings. Results show that the semantic search consistently returns a larger number of results than the keyword search and performs particularly well for place based, misspelled, obscure, or complex queries. While the semantic search does not capture all keyword based results, the datasets returned are overwhelmingly semantically similar, with high cosine similarity scores despite lower exact overlap. Rankings of the most relevant results differ substantially between tools, reflecting contrasting prioritisation strategies. Case studies demonstrate that the LLM based tool is robust to spelling errors, interprets geographic and contextual relevance effectively, and supports natural-language queries that keyword search fails to resolve. Overall, the findings suggest that LLM driven semantic search offers a substantial improvement for data discovery, complementing rather than fully replacing traditional keyword-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。