arXiv:2510.22521cs.CVcs.AI2025-10被引 6

让AI生成图片时更讲事实,避免编造细节。

Open Multimodal Retrieval-Augmented Factual Image Generation

  • 用动态网页检索+逐步优化提示词,让生成图像更符合真实信息。
  • 在十类场景下测试,事实一致性显著优于现有方法。
  • 适合需要高准确度图像生成的研究与应用,如新闻、教育。

大型多模态模型在生成逼真且符合提示的图像方面取得了显著进展,但常出现与可验证知识相悖的情况,尤其在涉及细粒度属性或时效性事件时。传统检索增强方法依赖静态数据源和浅层证据融合,难以实现对动态演变知识的有效支撑。为此,我们提出ORIG——一种面向事实性图像生成(FIG)的智能体式开放多模态检索增强框架。该框架通过迭代从网络中检索并筛选多模态证据,逐步将提炼后的知识融入提示词,引导图像生成。为支持系统评估,我们构建了涵盖感知、构图和时间维度共十个类别的FIG-Eval基准。实验表明,ORIG在事实一致性与整体图像质量上均显著超越强基线模型,凸显了开放多模态检索在事实性图像生成中的潜力。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have achieved remarkable progress in generating photorealistic and prompt-aligned images, but they often produce outputs that contradict verifiable knowledge, especially when prompts involve fine-grained attributes or time-sensitive events. Conventional retrieval-augmented approaches attempt to address this issue by introducing external information, yet they are fundamentally incapable of grounding generation in accurate and evolving knowledge due to their reliance on static sources and shallow evidence integration. To bridge this gap, we introduce ORIG, an agentic open multimodal retrieval-augmented framework for Factual Image Generation (FIG), a new task that requires both visual realism and factual grounding. ORIG iteratively retrieves and filters multimodal evidence from the web and incrementally integrates the refined knowledge into enriched prompts to guide generation. To support systematic evaluation, we build FIG-Eval, a benchmark spanning ten categories across perceptual, compositional, and temporal dimensions. Experiments demonstrate that ORIG substantially improves factual consistency and overall image quality over strong baselines, highlighting the potential of open multimodal retrieval for factual image generation.

图像生成事实性多模态检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。