用信息增益评估视觉证据价值,提升多模态生成的实用性。
Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation

- 从信息论角度定义视觉证据效用,以输出分布变化衡量价值。
- 在多个模型上显著超越现有基线,计算开销降低明显。
- 无需训练、轻量高效,适合追求效率与准确性的多模态系统。
视觉证据选择是多模态检索增强生成(RAG)的关键环节,但现有方法多依赖语义相关性或表层相似性,常与下游推理的实际效用不一致。本文从信息论视角重新定义证据效用为模型输出分布的信息增益。为克服答案空间优化的不可行性,引入潜在的证据帮助度概念,并在温和假设下证明:对潜在变量的信息增益排序等价于答案空间的效用排序。进一步提出一种无训练、代理加速的框架,利用轻量级多模态模型高效估算证据效用。在MRAG-Bench和Visual-RAG上的实验表明,该方法在多种模型家族中持续优于当前最优基线,同时实现显著的计算成本降低。
原文摘要 · Abstract (English)
Visual evidence selection is a critical component of multimodal retrieval-augmented generation (RAG), yet existing methods typically rely on semantic relevance or surface-level similarity, which are often misaligned with the actual utility of visual evidence for downstream reasoning. We reformulate multimodal evidence selection from an information-theoretic perspective by defining evidence utility as the information gain induced on a model's output distribution. To overcome the intractability of answer-space optimization, we introduce a latent notion of evidence helpfulness and theoretically show that, under mild assumptions, ranking evidence by information gain on this latent variable is equivalent to answer-space utility. We further propose a training-free, surrogate-accelerated framework that efficiently estimates evidence utility using lightweight multimodal models. Experiments on MRAG-Bench and Visual-RAG across multiple model families demonstrate that our method consistently outperforms state-of-the-art RAG baselines while achieving substantial reductions in computational cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。