用大模型解析文档,让RAG系统更好理解复杂多模态文件
Advanced ingestion process powered by LLM parsing for RAG system
- 用大模型驱动的OCR和节点化提取,打通文档内容关联
- 在多个知识库测试中,答案相关性和信息准确性均提升
- 适合需要处理扫描件、报告等复杂文档的RAG应用
检索增强生成(RAG)系统在处理结构复杂多样的多模态文档时面临挑战。本文提出一种基于大语言模型(LLM)驱动的OCR的多策略解析方法,可从演示文稿及高文本密度文件(无论是否扫描)中提取内容。该方法采用基于节点的提取技术,建立不同信息类型间的关联,并生成上下文感知的元数据。通过引入多模态组装代理(Multimodal Assembler Agent)与灵活嵌入策略,系统显著提升了文档理解和检索能力。在多个知识库上的实验验证了其有效性,结果显示答案相关性与信息忠实度均有提升。
原文摘要 · Abstract (English)
Retrieval Augmented Generation (RAG) systems struggle with processing multimodal documents of varying structural complexity. This paper introduces a novel multi-strategy parsing approach using LLM-powered OCR to extract content from diverse document types, including presentations and high text density files both scanned or not. The methodology employs a node-based extraction technique that creates relationships between different information types and generates context-aware metadata. By implementing a Multimodal Assembler Agent and a flexible embedding strategy, the system enhances document comprehension and retrieval capabilities. Experimental evaluations across multiple knowledge bases demonstrate the approach's effectiveness, showing improvements in answer relevancy and information faithfulness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。