自动生成图文并茂的维基风格文章,提升内容准确性和视觉吸引力。
WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
- 引入多视角自我反思机制,提升信息可靠性和文章连贯性。
- 在新基准WikiSeek上性能优于旧方法8%-29%,生成更准确的多模态内容。
- 适合对知识生成、多模态内容创作感兴趣的开发者与研究者。
知识发现与收集是高度依赖智力的任务,传统上需要大量人工投入以保证高质量输出。近期研究探索了基于多智能体框架的自动化维基风格文章生成,通过从互联网检索和整合信息实现。然而,这些方法主要聚焦于纯文本生成,忽视了多模态内容在增强信息丰富性和用户参与度方面的重要性。本文提出WikiAutoGen,一个面向自动化多模态维基风格文章生成的新系统。不同于以往方法,WikiAutoGen不仅检索并整合相关图像,还同步丰富文本内容,提升文章深度与视觉表现力。为增强事实准确性与全面性,我们设计了一种多视角自我反思机制,从多个角度批判性评估检索内容,从而提高可靠性、广度与连贯性。此外,我们构建了WikiSeek基准,包含带有文本与图像表征的维基文章,用于评估在更具挑战性主题上的多模态知识生成能力。实验结果表明,WikiAutoGen在WikiSeek基准上相较现有方法提升8%-29%,生成更准确、连贯且视觉丰富的维基风格文章。代码与示例已公开于https://wikiautogen.github.io/。
原文摘要 · Abstract (English)
Knowledge discovery and collection are intelligence-intensive tasks that traditionally require significant human effort to ensure high-quality outputs. Recent research has explored multi-agent frameworks for automating Wikipedia-style article generation by retrieving and synthesizing information from the internet. However, these methods primarily focus on text-only generation, overlooking the importance of multimodal content in enhancing informativeness and engagement. In this work, we introduce WikiAutoGen, a novel system for automated multimodal Wikipedia-style article generation. Unlike prior approaches, WikiAutoGen retrieves and integrates relevant images alongside text, enriching both the depth and visual appeal of generated content. To further improve factual accuracy and comprehensiveness, we propose a multi-perspective self-reflection mechanism, which critically assesses retrieved content from diverse viewpoints to enhance reliability, breadth, and coherence, etc. Additionally, we introduce WikiSeek, a benchmark comprising Wikipedia articles with topics paired with both textual and image-based representations, designed to evaluate multimodal knowledge generation on more challenging topics. Experimental results show that WikiAutoGen outperforms previous methods by 8%-29% on our WikiSeek benchmark, producing more accurate, coherent, and visually enriched Wikipedia-style articles. Our code and examples are available at https://wikiautogen.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。