将论文转化为可交互的网页系统,让技术内容动态可操作。
PaperVoyager : Building Interactive Web with Visual Language Models
- 从论文自动构建可执行的交互式网页系统
- 在19篇论文上验证,生成效果显著优于现有方法
- 适合科研人员快速理解复杂机制与状态变化
视觉语言模型的进步使得复杂推理、工具使用和文档理解成为可能。然而,现有文档代理主要将论文转换为静态摘要、网页或幻灯片,难以呈现涉及动态机制与状态转移的技术论文。本文提出一种端到端的论文转交互系统代理,给定一篇PDF论文,无需人工干预即可完成论文理解、系统建模和交互网页合成,使用户能操作输入并观察动态行为。为此,我们构建了一个包含19篇论文及其专家制作的交互系统作为真实标签的基准。进一步提出PaperVoyager框架,通过显式建模机制与交互逻辑提升生成质量。实验表明,PaperVoyager显著提升了生成系统的质量,为交互式科学论文理解提供了新范式。
原文摘要 · Abstract (English)
Recent advances in visual language models have enabled autonomous agents for complex reasoning, tool use, and document understanding. However, existing document agents mainly transform papers into static artifacts such as summaries, webpages, or slides, which are insufficient for technical papers involving dynamic mechanisms and state transitions. In this work, we propose a Paper-to-Interactive-System Agent that converts research papers into executable interactive web systems. Given a PDF paper, the agent performs end-to-end processing without human intervention, including paper understanding, system modeling, and interactive webpage synthesis, enabling users to manipulate inputs and observe dynamic behaviors. To evaluate this task, we introduce a benchmark of 19 research papers paired with expert-built interactive systems as ground truth. We further propose PaperVoyager, a structured generation framework that explicitly models mechanisms and interaction logic during synthesis. Experiments show that PaperVoyager significantly improves the quality of generated interactive systems, offering a new paradigm for interactive scientific paper understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。