arXiv:2410.20163cs.IRcs.CL2024-10NAACL被引 12

统一处理异构知识与指令,提升复杂检索场景下的准确性。

UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers

  • 构建统一检索空间,融合多种类型知识与指令
  • 在1000万条数据上达54.23%相对提升
  • 适合开放域异构问答系统开发者使用

现有信息检索模型通常假设知识源和用户查询具有同质结构,限制了其在真实场景中的应用。本文提出UniHGKR,一种统一的指令感知异构知识检索器,能够建立统一的检索空间,并根据多样化用户指令检索指定类型的知识。该框架包含三个阶段:异构自监督预训练、文本锚定嵌入对齐和指令感知微调,具备跨不同检索场景的泛化能力。支持基于BERT的版本及在大语言模型上训练的UniHGKR-7B版本。我们还提出了首个原生异构知识检索基准CompMix-IR,涵盖两种检索场景,包含超过9,400个问答对和1000万条语料,覆盖四类数据类型。大量实验表明,UniHGKR在CompMix-IR上持续优于现有方法,两个场景分别实现6.36%和54.23%的相对提升。此外,将其用于开放域异构问答系统,在流行ConvMix任务上达到新最优,绝对提升达5.90分。

原文摘要 · Abstract (English)

Existing information retrieval (IR) models often assume a homogeneous structure for knowledge sources and user queries, limiting their applicability in real-world settings where retrieval is inherently heterogeneous and diverse. In this paper, we introduce UniHGKR, a unified instruction-aware heterogeneous knowledge retriever that (1) builds a unified retrieval space for heterogeneous knowledge and (2) follows diverse user instructions to retrieve knowledge of specified types. UniHGKR consists of three principal stages: heterogeneous self-supervised pretraining, text-anchored embedding alignment, and instruction-aware retriever fine-tuning, enabling it to generalize across varied retrieval contexts. This framework is highly scalable, with a BERT-based version and a UniHGKR-7B version trained on large language models. Also, we introduce CompMix-IR, the first native heterogeneous knowledge retrieval benchmark. It includes two retrieval scenarios with various instructions, over 9,400 question-answer (QA) pairs, and a corpus of 10 million entries, covering four different types of data. Extensive experiments show that UniHGKR consistently outperforms state-of-the-art methods on CompMix-IR, achieving up to 6.36% and 54.23% relative improvements in two scenarios, respectively. Finally, by equipping our retriever for open-domain heterogeneous QA systems, we achieve a new state-of-the-art result on the popular ConvMix task, with an absolute improvement of up to 5.90 points.

知识检索异构数据指令感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。