arXiv:2501.07024cs.AIcs.IR2025-01被引 1

用大模型让档案系统能懂人话,搜得准、找得快。

A Proposed Large Language Model-Based Smart Search for Archive System

  • 基于RAG框架,把自然语言查询和非文本数据转为可检索文本。
  • 多语言查询下准确率显著提升,组件实验验证有效性。
  • 适合需要智能检索的档案馆、图书馆等机构使用。

本研究提出一种基于大语言模型(LLM)的智能档案搜索框架,利用检索增强生成(RAG)技术,提升数字档案系统的信息检索能力。该系统通过自然语言查询处理与非文本数据的语义化转换,结合先进的元数据生成、混合检索机制、路由查询引擎及响应合成模块,在四项实验中评估了大模型效率、混合检索优化、多语言查询处理以及各组件影响。结果表明,相较于传统方法,系统在搜索精度与相关性上均有显著提升,展现出人工智能驱动系统在现代档案管理中的变革潜力。

原文摘要 · Abstract (English)

This study presents a novel framework for smart search in digital archival systems, leveraging the capabilities of Large Language Models (LLMs) to enhance information retrieval. By employing a Retrieval-Augmented Generation (RAG) approach, the framework enables the processing of natural language queries and transforming non-textual data into meaningful textual representations. The system integrates advanced metadata generation techniques, a hybrid retrieval mechanism, a router query engine, and robust response synthesis, the results proved search precision and relevance. We present the architecture and implementation of the system and evaluate its performance in four experiments concerning LLM efficiency, hybrid retrieval optimizations, multilingual query handling, and the impacts of individual components. Obtained results show significant improvements over conventional approaches and have demonstrated the potential of AI-powered systems to transform modern archival practices.

智能搜索大模型档案系统RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。