arXiv:2509.10697cs.CL2025-09KDD综述被引 26

解决大模型幻觉与知识滞后问题,通过检索与结构化增强生成能力

A Survey on Retrieval And Structuring Augmented Generation with Large Language Models

  • 融合外部检索与结构化知识,提升生成准确性
  • 涵盖稀疏、稠密、混合检索及分类、抽取等结构化技术
  • 适合关注大模型可靠性的研究者与应用开发者

大型语言模型(LLMs)在文本生成与推理方面展现出卓越能力,但在实际应用中面临幻觉生成、知识过时和领域专长有限等挑战。检索与结构化(RAS)增强生成通过动态信息检索与结构化知识表示,缓解上述问题。本综述系统分析了三方面内容:(1) 外部知识访问的稀疏、稠密及混合检索机制;(2) 税收构建、层次分类与信息抽取等文本结构化技术;(3) 结构化表示通过提示工程、推理框架与知识嵌入与大模型的融合方式。同时指出了检索效率、结构质量与知识整合中的技术挑战,并展望了多模态检索、跨语言结构与交互系统等研究机遇。本文为研究人员与实践者提供了RAS方法、应用与未来方向的全面洞察。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have revolutionized natural language processing with their remarkable capabilities in text generation and reasoning. However, these models face critical challenges when deployed in real-world applications, including hallucination generation, outdated knowledge, and limited domain expertise. Retrieval And Structuring (RAS) Augmented Generation addresses these limitations by integrating dynamic information retrieval with structured knowledge representations. This survey (1) examines retrieval mechanisms including sparse, dense, and hybrid approaches for accessing external knowledge; (2) explore text structuring techniques such as taxonomy construction, hierarchical classification, and information extraction that transform unstructured text into organized representations; and (3) investigate how these structured representations integrate with LLMs through prompt-based methods, reasoning frameworks, and knowledge embedding techniques. It also identifies technical challenges in retrieval efficiency, structure quality, and knowledge integration, while highlighting research opportunities in multimodal retrieval, cross-lingual structures, and interactive systems. This comprehensive overview provides researchers and practitioners with insights into RAS methods, applications, and future directions.

大模型检索增强知识结构化综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。