用大模型增强检索,提升学术术语匹配准确率
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
- 融合大模型语义理解能力与相似度矩阵加权策略
- 在竞赛数据集上取得0.20726的最终得分(第二名)
- 适合需要精准学术知识检索的研究者参考
在技术迅猛发展和信息快速更新的时代,为研究者和公众提供跨领域高质量、前沿的学术见解已成为迫切需求。KDD Cup 2024 AQA挑战旨在推动检索模型的发展,以从相关论文中识别出适用于科学问题的学术术语。本文介绍由Robo Space提出的LLM-KnowSimFuser方法,在该比赛中获得第二名。基于大模型在多项任务中的优异表现,我们通过对提供的数据集进行细致分析,首先利用大模型增强的预训练检索模型进行微调与推理,将大模型强大的语言理解能力和开放域知识引入该任务;随后基于推理结果生成的相似度矩阵,实施加权融合策略。最终在比赛数据集上的实验表明,所提方法具有显著优势,取得了0.20726的最终排行榜分数。
原文摘要 · Abstract (English)
In an era marked by robust technological growth and swift information renewal, furnishing researchers and the populace with top-tier, avant-garde academic insights spanning various domains has become an urgent necessity. The KDD Cup 2024 AQA Challenge is geared towards advancing retrieval models to identify pertinent academic terminologies from suitable papers for scientific inquiries. This paper introduces the LLM-KnowSimFuser proposed by Robo Space, which wins the 2nd place in the competition. With inspirations drawed from the superior performance of LLMs on multiple tasks, after careful analysis of the provided datasets, we firstly perform fine-tuning and inference using LLM-enhanced pre-trained retrieval models to introduce the tremendous language understanding and open-domain knowledge of LLMs into this task, followed by a weighted fusion based on the similarity matrix derived from the inference results. Finally, experiments conducted on the competition datasets show the superiority of our proposal, which achieved a score of 0.20726 on the final leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。