用多智能体系统自动优化科学论文,保持内容与排版一致。
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
- 六类智能体协作分析文档布局、内容与指令
- 在多个任务上达到86.7%语义一致性、93.9%排版保真度
- 适合需要精准编辑学术论文的研究者使用
科学文献以PDF形式爆炸式增长,亟需高效准确的文档理解、摘要与内容优化工具。传统方法难以处理复杂版式与多模态内容,直接使用大语言模型(LLMs)和视觉-语言大模型(LVLMs)又缺乏精确控制。本文提出DocRefine,一个基于自然语言指令的智能框架,用于科学PDF文档的智能理解、内容优化与自动生成。该框架利用先进LVLM(如GPT-4o),通过六个专业化协作智能体构成的多智能体系统:版式与结构分析、多模态内容理解、指令分解、内容优化、摘要生成与一致性验证,形成闭环反馈架构,确保语义准确性与视觉保真度。在综合性评测集DocEditBench上,其整体表现优于现有最优基线,分别取得86.7%的语义一致性得分(SCS)、93.9%的排版保真度指数(LFI)和85.0%的指令遵循率(IAR)。结果表明,DocRefine在处理复杂多模态文档编辑任务中具有卓越能力,有效保留语义完整性并维持视觉一致性,显著推进自动化科学文档处理技术的发展。
原文摘要 · Abstract (English)
The exponential growth of scientific literature in PDF format necessitates advanced tools for efficient and accurate document understanding, summarization, and content optimization. Traditional methods fall short in handling complex layouts and multimodal content, while direct application of Large Language Models (LLMs) and Vision-Language Large Models (LVLMs) lacks precision and control for intricate editing tasks. This paper introduces DocRefine, an innovative framework designed for intelligent understanding, content refinement, and automated summarization of scientific PDF documents, driven by natural language instructions. DocRefine leverages the power of advanced LVLMs (e.g., GPT-4o) by orchestrating a sophisticated multi-agent system comprising six specialized and collaborative agents: Layout & Structure Analysis, Multimodal Content Understanding, Instruction Decomposition, Content Refinement, Summarization & Generation, and Fidelity & Consistency Verification. This closed-loop feedback architecture ensures high semantic accuracy and visual fidelity. Evaluated on the comprehensive DocEditBench dataset, DocRefine consistently outperforms state-of-the-art baselines across various tasks, achieving overall scores of 86.7% for Semantic Consistency Score (SCS), 93.9% for Layout Fidelity Index (LFI), and 85.0% for Instruction Adherence Rate (IAR). These results demonstrate DocRefine's superior capability in handling complex multimodal document editing, preserving semantic integrity, and maintaining visual consistency, marking a significant advancement in automated scientific document processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。