arXiv:2504.05634cs.DBcs.IR2025-04被引 8

用轻量语言模型实现跨异构数据库的统一语义查询

Simplifying Data Integration: SLM-Driven Systems for Unified Semantic Queries Across Heterogeneous Databases

  • 用小模型结合检索增强生成,自动提取结构化数据
  • 在多实体问答任务中准确率显著提升,响应速度更快
  • 适合资源受限场景,无需领域知识也能部署

将异构数据库整合为统一查询框架仍面临挑战,尤其在资源受限环境下。本文提出一种基于小型语言模型(SLM)的系统,融合轻量级检索增强生成(RAG)与语义感知数据建模技术,实现跨多种数据格式的高效、准确、可扩展查询。通过集成MiniRAG的语义感知异构图索引与拓扑增强检索,结合SLM驱动的结构化数据提取,有效解决了传统方法在多实体问答和复杂语义查询中的不足。实验表明,该系统在准确率和效率上均表现优异;引入的语义熵作为无监督评估指标,可有效反映模型不确定性。本工作开创了一种低成本、领域无关的下一代数据库系统解决方案。

原文摘要 · Abstract (English)

The integration of heterogeneous databases into a unified querying framework remains a critical challenge, particularly in resource-constrained environments. This paper presents a novel Small Language Model(SLM)-driven system that synergizes advancements in lightweight Retrieval-Augmented Generation (RAG) and semantic-aware data structuring to enable efficient, accurate, and scalable query resolution across diverse data formats. By integrating MiniRAG's semantic-aware heterogeneous graph indexing and topology-enhanced retrieval with SLM-powered structured data extraction, our system addresses the limitations of traditional methods in handling Multi-Entity Question Answering (Multi-Entity QA) and complex semantic queries. Experimental results demonstrate superior performance in accuracy and efficiency, while the introduction of semantic entropy as an unsupervised evaluation metric provides robust insights into model uncertainty. This work pioneers a cost-effective, domain-agnostic solution for next-generation database systems.

数据库小模型语义查询RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。