让检索模型理解问题背后的因果关系,提升专业领域问答准确率。
Causal Retrieval with Semantic Consideration
- 用语义+因果双重目标训练检索模型,捕捉深层关系
- 在大规模检索中显著优于现有模型,尤其擅长处理因果类问题
- 无需微调即可跨科学领域通用,适合医疗法律等高精度场景
大型语言模型(LLMs)的进展显著提升了对话AI的表现。为拓展其在生物医学、法律等知识密集型领域的应用,常结合信息检索(IR)系统,基于文档生成回答。然而,为支持此类应用,IR系统需超越表面语义匹配,准确捕捉查询意图中的因果关系。现有模型主要依赖表面语义相似性,忽视因果等深层结构。为此,我们提出CAWAI,一个通过语义与因果双重目标训练的检索模型。大量实验表明,CAWAI在多种因果检索任务中表现优异,尤其在大规模检索设置下优势明显。同时,该模型在科学领域问答任务中展现出强零样本泛化能力。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have significantly enhanced the performance of conversational AI systems. To extend their capabilities to knowledge-intensive domains such as biomedical and legal fields, where the accuracy is critical, LLMs are often combined with information retrieval (IR) systems to generate responses based on retrieved documents. However, for IR systems to effectively support such applications, they must go beyond simple semantic matching and accurately capture diverse query intents, including causal relationships. Existing IR models primarily focus on retrieving documents based on surface-level semantic similarity, overlooking deeper relational structures such as causality. To address this, we propose CAWAI, a retrieval model that is trained with dual objectives: semantic and causal relations. Our extensive experiments demonstrate that CAWAI outperforms various models on diverse causal retrieval tasks especially under large-scale retrieval settings. We also show that CAWAI exhibits strong zero-shot generalization across scientific domain QA tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。