用大模型理解语义,让科研数据搜索更智能
From keywords to semantics: Perceptions of large language models in data discovery
- 通过大模型理解自然语言查询,无需精确关键词匹配
- 研究发现透明性功能能有效克服用户对大模型的接受障碍
- 适合关注数据发现系统设计的研究者和开发者
当前数据发现方法依赖元数据与查询之间的关键词匹配,要求研究者掌握他人使用的精确术语,导致易遗漏相关数据。大语言模型(LLMs)可通过自然语言理解消除此需求。我们基于人中心人工智能(HCAI)方法,开展焦点小组访谈(N=27),探究研究人员对使用大模型进行数据发现的看法。研究构建的概念模型表明,尽管潜在优势明显,但现有障碍仍阻碍研究人员采用大模型替代现有技术。然而,增强透明性特征可有效克服这些障碍。该模型可指导开发者设计提升接受度的功能。
原文摘要 · Abstract (English)
Current approaches to data discovery match keywords between metadata and queries. This matching requires researchers to know the exact wording that other researchers previously used, creating a challenging process that could lead to missing relevant data. Large Language Models (LLMs) could enhance data discovery by removing this requirement and allowing researchers to ask questions with natural language. However, we do not currently know if researchers would accept LLMs for data discovery. Using a human-centered artificial intelligence (HCAI) focus, we ran focus groups (N = 27) to understand researchers' perspectives towards LLMs for data discovery. Our conceptual model shows that the potential benefits are not enough for researchers to use LLMs instead of current technology. Barriers prevent researchers from fully accepting LLMs, but features around transparency could overcome them. Using our model will allow developers to incorporate features that result in an increased acceptance of LLMs for data discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。