arXiv:2512.15365cs.DBcs.IR2025-12

用大模型理解自然语言,智能搜索多组学元数据。

ArcBERT: An LLM-based Search Engine for Exploring Integrated Multi-Omics Metadata

  • 基于大模型实现自然语言查询,无需关键词输入。
  • 支持元数据结构与层级理解,适应多样查询方式。
  • 适合生物医学研究者快速探索复杂数据集。

科研数据管理生态系统中的传统搜索应用对发现和探索研究数据的结构化元数据至关重要。通常,文本搜索引擎要求用户提交基于关键词的查询,而非自然语言。然而,利用在领域特定内容上训练的大语言模型(LLM)进行专业自然语言处理任务正日益普遍。我们提出 ArcBERT,一个基于大语言模型的系统,专为整合的元数据探索而设计。ArcBERT 能理解自然语言查询,并依赖语义匹配,不同于传统搜索应用。值得注意的是,ArcBERT 还能理解元数据内部的结构与层级,从而有效应对多样的用户查询模式。

原文摘要 · Abstract (English)

Traditional search applications within Research Data Management (RDM) ecosystems are crucial in helping users discover and explore the structured metadata from the research datasets. Typically, text search engines require users to submit keyword-based queries rather than using natural language. However, using Large Language Models (LLMs) trained on domain-specific content for specialized natural language processing (NLP) tasks is becoming increasingly common. We present ArcBERT, an LLM-based system designed for integrated metadata exploration. ArcBERT understands natural language queries and relies on semantic matching, unlike traditional search applications. Notably, ArcBERT also understands the structure and hierarchies within the metadata, enabling it to handle diverse user querying patterns effectively.

大模型元数据搜索多组学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。