arXiv:2503.23229cs.LG2025-03被引 4

用arXiv论文库自动生成引言中参考文献分析,避免AI虚构引用。

Citegeist: Automated Generation of Related Work Analysis on the arXiv Corpus

  • 基于arXiv文档的动态检索增强生成,确保引用真实可靠。
  • 通过多阶段过滤与摘要匹配,精准生成相关工作段落。
  • 适合科研人员快速撰写论文引言,支持多种大模型部署。

大型语言模型为高质量文本生成带来新机遇,但其在学术界的应用受限于虚构无效引用及缺乏对相关科学文献知识库的直接访问。本文提出Citegeist:一种基于arXiv语料库的动态检索增强生成(RAG)应用流程,可自动生成相关工作章节及其他基于引用的输出内容。为此,我们结合嵌入相似性匹配、摘要生成与多阶段过滤机制。为适应文献库持续增长,还提出一种优化的新旧论文融合方式。为便于科研社区使用,我们发布了网站(https://citegeist.org)及兼容多种LLM实现的集成工具包。

原文摘要 · Abstract (English)

Large Language Models provide significant new opportunities for the generation of high-quality written works. However, their employment in the research community is inhibited by their tendency to hallucinate invalid sources and lack of direct access to a knowledge base of relevant scientific articles. In this work, we present Citegeist: An application pipeline using dynamic Retrieval Augmented Generation (RAG) on the arXiv Corpus to generate a related work section and other citation-backed outputs. For this purpose, we employ a mixture of embedding-based similarity matching, summarization, and multi-stage filtering. To adapt to the continuous growth of the document base, we also present an optimized way of incorporating new and modified papers. To enable easy utilization in the scientific community, we release both, a website (https://citegeist.org), as well as an implementation harness that works with several different LLM implementations.

文献生成RAGarXiv引言写作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。