arXiv:2502.16868cs.DBcs.AI2025-02被引 4

用图结构自动分析文献,一键生成综述报告。

Graphy'our Data: Towards End-to-End Modeling, Exploring and Generating Report from Raw Data

  • 将原始论文转为事实与维度节点构成的图结构
  • 支持迭代探索,自动生成高质量综述报告
  • 适合科研人员快速完成文献调研

大型语言模型在检索增强生成和自主AI代理任务中表现优异,但在处理需逐步探索、分析与整合的大量非结构化文档(如文献综述)时仍显不足。为此,我们提出一种名为Graphy的端到端平台,解决‘渐进式文档调查’挑战。Graphy包含离线的Scrapper模块,可将原始文档转化为包含超过5万篇论文及其引用关系的结构化图;以及在线的Surveyor模块,支持迭代探索与基于LLM的报告生成。该平台显著提升了文献调研效率,演示视频见 https://youtu.be/uM4nzkAdGlM。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have recently demonstrated remarkable performance in tasks such as Retrieval-Augmented Generation (RAG) and autonomous AI agent workflows. Yet, when faced with large sets of unstructured documents requiring progressive exploration, analysis, and synthesis, such as conducting literature survey, existing approaches often fall short. We address this challenge -- termed Progressive Document Investigation -- by introducing Graphy, an end-to-end platform that automates data modeling, exploration and high-quality report generation in a user-friendly manner. Graphy comprises an offline Scrapper that transforms raw documents into a structured graph of Fact and Dimension nodes, and an online Surveyor that enables iterative exploration and LLM-driven report generation. We showcase a pre-scrapped graph of over 50,000 papers -- complete with their references -- demonstrating how Graphy facilitates the literature-survey scenario. The demonstration video can be found at https://youtu.be/uM4nzkAdGlM.

文献综述知识图谱LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。