Wyvern让生成的多模态报告图文并茂且有据可查,适合技术文档自动化。
Wyvern: An Agentic Framework for Generating Grounded Multimodal Reports

- 多智能体协作生成带图表与引用的报告
- 87%情况下图表信息量优于基线,报告实用性提升63%-100%
- 引用召回率最高提升2.3倍,精度提升1.6倍
在人工智能驱动创新的时代,知识增长速度加快,难以跟上。尽管生成模型被广泛用于内容合成,但常缺乏信息根基。为此,我们提出 Wyvern——一种用于自动生成具备信息根基的多模态技术报告的多智能体框架。Wyvern 支持生成融合图像、表格和文本的多模态输出,并附带参考文献。特别注重内容的可验证性,引入了声明自动修订阶段。通过人工评估,其生成的图表在87%的情况下被认为比近期基线更富信息量;报告实用性在63%至100%的案例中优于三种替代方法。自动评估显示,与基线相比,Wyvern 的引用召回率最高提升2.3倍,引用精确率提升1.6倍。
原文摘要 · Abstract (English)
In the current artificial intelligence-driven innovation era, the pace of knowledge growth is accelerating, and is hard to keep up with. While generative models are increasingly used to synthesize content, they often lack in information grounding. To address these peculiarities of our time, we propose Wyvern, a multi-agent framework for the automated generation of grounded, multimodal technical reports. Wyvern allows for the generation of multimodal outputs, integrating images, tables, and text with supporting references in a unified report. Additionally, a particular focus is placed on the grounding of the content, with the implementation of a claims auto-revision stage. We conduct a human evaluation study to assess the quality of our proposed framework. The results show that the figures' informativeness is perceived as superior to that of a recent baseline in 87% of cases. Furthermore, Wyvern's reports are rated as more useful than those produced by three alternative methods in 63% to 100% of instances. We also carry out automatic evaluations showing that Wyvern gains up to 2.3$\times$ in citation recall and 1.6$\times$ in citation precision with respect to the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。